Pilot 3 rows before a Sume bulk batch and estimate spend from receipts
Run three rows first, read usage.debited_usd_micros from each receipt, and scale up. Skip billable_amount_usd_micros: it leaves out the agent turn.

How do I estimate what a 100-row batch will cost?
Run a small queue of three representative rows, wait for every item to be terminal, and read the cost from the receipts. Then multiply the average by the real row count and add headroom. Pick rows that include your longest input, since spend follows what the agent generates for that row.
Three is a judgment call, not a rule. It is enough to catch a Format that spends ten times what you expected, and cheap enough to throw away.
Which usage field is the real cost?
Each run receipt has a usage object, and it can be null while the run is in flight. Two fields look similar. usage.billable_amount_usd_micros is generation spend only and excludes the agent's own LLM turn. usage.debited_usd_micros is the amount actually taken from your wallet. For a budget, use the debited figure. A micro is one millionth of a dollar, so 1,800,000 micros is $1.80.
Read them only from terminal runs, and skip items whose run_id is null.
def estimate(receipts, total_rows, headroom=1.25):
debited = [r["usage"]["debited_usd_micros"]
for r in receipts if r.get("usage")]
if not debited:
raise ValueError("no usage on any receipt yet")
avg = sum(debited) / len(debited)
return round(avg * total_rows * headroom / 1_000_000, 2)
pilot = [{"usage": {"debited_usd_micros": m}}
for m in (1_800_000, 2_100_000, 2_400_000)]
print(estimate(pilot, 100)) # dollars
What does that example give, and what does it miss?
The three sample receipts average 2,100,000 micros, or $2.10. At 100 rows with 25 percent headroom the sample prints 262.5, meaning $262.50. Those numbers are an illustration of the arithmetic, not a price list for any Format.
The estimate is an average, so a heavy row can cost more. Put a per-item generation_spend_cap_usd on the real batch to bound each row, and remember there is no queue-wide cap.
- Pilot with the same Format version you will run; a mid-batch edit changes the package.
- Use the same
modelper item as the real run, since it picks the orchestrator. - Re-estimate when the input gets longer or the Format instruction changes.
Sources
Related posts
More in Formats
- Retry one scene and keep the voice track: Format previous_run_id
A new voice is not a reason to redo a whole video. Continue a finished Format run with previous_run_id and an instruction that touches one scene only.
- Spreadsheet with more than 64 columns: nest them in one Sume input key
A Sume run input allows 64 top-level keys and 2 MiB. Nested keys do not count toward the 64, so a 90-column sheet row fits under a single key.
- Stage a Sume Format change with no branches: second slug, If-Match
Sume Formats have no branches or revert. Test a change in a second Format slug, then promote it to the live one with a single If-Match PUT of the files.
- Stocking stuffer ads for 100 products: one bulk run, 16 at a time
One Sume bulk run takes 1-100 items and 1-16 parallel workers. Here is how to set it up for a stocking stuffer catalog and check the failed count at the end.
Written by Sume