Pilot 3 rows before a Sume bulk batch and estimate spend from receipts

Run three rows first, read usage.debited_usd_micros from each receipt, and scale up. Skip billable_amount_usd_micros: it leaves out the agent turn.

3 min readSume
All posts

How do I estimate what a 100-row batch will cost?

Run a small queue of three representative rows, wait for every item to be terminal, and read the cost from the receipts. Then multiply the average by the real row count and add headroom. Pick rows that include your longest input, since spend follows what the agent generates for that row.

Three is a judgment call, not a rule. It is enough to catch a Format that spends ten times what you expected, and cheap enough to throw away.

Which usage field is the real cost?

Each run receipt has a usage object, and it can be null while the run is in flight. Two fields look similar. usage.billable_amount_usd_micros is generation spend only and excludes the agent's own LLM turn. usage.debited_usd_micros is the amount actually taken from your wallet. For a budget, use the debited figure. A micro is one millionth of a dollar, so 1,800,000 micros is $1.80.

Read them only from terminal runs, and skip items whose run_id is null.

def estimate(receipts, total_rows, headroom=1.25):
    debited = [r["usage"]["debited_usd_micros"]
               for r in receipts if r.get("usage")]
    if not debited:
        raise ValueError("no usage on any receipt yet")
    avg = sum(debited) / len(debited)
    return round(avg * total_rows * headroom / 1_000_000, 2)

pilot = [{"usage": {"debited_usd_micros": m}}
         for m in (1_800_000, 2_100_000, 2_400_000)]
print(estimate(pilot, 100))  # dollars

What does that example give, and what does it miss?

The three sample receipts average 2,100,000 micros, or $2.10. At 100 rows with 25 percent headroom the sample prints 262.5, meaning $262.50. Those numbers are an illustration of the arithmetic, not a price list for any Format.

The estimate is an average, so a heavy row can cost more. Put a per-item generation_spend_cap_usd on the real batch to bound each row, and remember there is no queue-wide cap.

  • Pilot with the same Format version you will run; a mid-batch edit changes the package.
  • Use the same model per item as the real run, since it picks the orchestrator.
  • Re-estimate when the input gets longer or the Format instruction changes.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume