Bulk queue worst case: 100 items x per-item cap, $300 not $40,000
Each child in a Sume bulk queue is a run with its own spend cap. 100 items at $3 cap at most $300; the default $400 cap would allow $40,000. Set caps per item.

A Sume bulk queue has no queue-level spend cap. Each of up to 100 items is an ordinary Format run with its own cap, so the worst case for generation spend is the number of items times the per-item cap. With 100 items at generation_spend_cap_usd: 3, that is $300. If every item inherited a Format that never named a cap, the default is $400 per run, which makes the ceiling $40,000.
Where the cap lives
A bulk request is { concurrency, items }, with concurrency from 1 to 16 and 1 to 100 items. Each item has the same body as a single POST …/runs, including generation_spend_cap_usd. The queue itself has no cap field. A run can never spend more than its own effective cap, and the receipt shows it as usage.generation_spend_cap_usd_micros.
| Per-item cap | Items | Ceiling |
|---|---|---|
| $3 | 100 | $300 |
| $25 | 100 | $2,500 |
| $120 | 100 | $12,000 |
| Inherited $400 default | 100 | $40,000 |
null ($500 maximum) | 100 | $50,000 |
The rules for caps
The request cap can be a number up to 500. Sume accepts a number above the Format's own cap and does not clamp it. null uses the platform maximum, $500. A cap of 0 or above 500 is 400. Because Sume does not clamp, a number you set per item is the number that applies. Set it deliberately on every item rather than relying on what the Format carries.
What the cap does not cover
The cap counts generation spend: reserved and captured amounts for the media the run makes. It does not include the LLM turn of the orchestrating agent, so usage.billable_amount_usd_micros is not the total cost of the run. The wallet cost is usage.debited_usd_micros, which includes the LLM row. Plan the budget on the debit, and use the cap as a guard rail.
The wallet still gates each child. If a child cannot start because of wallet or admission, that item becomes failed with the create-run error, and the window moves to the next queued item. The create already returned 202.
Why per-item caps beat one big number
A cap on every item turns an unbounded failure into a bounded one. If a bad prompt makes one item loop on retries, that item stops at its cap and the other 99 are untouched. A single high number applied to all items protects none of them. The bulk controller also runs every item with on_active_run: "allow", so items do not wait for each other beyond the concurrency window.
A sizing method
Run one item first as a single run with a generous cap and read usage on the receipt. Take the spend, add a margin of your own, and put that number on every item of the bulk queue. If a typical item takes $1.20, a $3 cap gives room for a retry-heavy item without letting any one item run away.
Then multiply: 100 x $3 = $300 is your ceiling, and 100 x $1.20 = $120 is your expectation. If the ceiling is too high for your wallet, split the sheet into two queues of 50 with different idempotency keys. Mint a fresh Idempotency-Key per batch, since a replayed key returns the old queue.
Sources
Related posts
More in Formats
- Cancel a Sume bulk queue: there is no queue cancel, so cancel children
Sume's API has no cancel-queue endpoint. Cancel each child with POST /v1/format-runs/{run_id}/cancel; the item frees its slot, and you pay for what already ran.
- Edited one ad hook and resent the bulk run? Same key gives 409
Reusing an Idempotency-Key with a changed Sume bulk-run payload returns 409 idempotency_conflict. Send only the edited hook under a new key.
- Format run incomplete_assembly: continue it, do not pay twice
incomplete_assembly means the run hit its time limit with generation jobs unfinished. Read pending_job_count, then continue with previous_run_id.
- Format run mcp_unavailable: failed before the model, charged false
mcp_unavailable means the per-turn tools never attached, so the run stopped before any turn. details.charged is false. Retry with a new Idempotency-Key.
Written by Sume