Run 300 ad variants as three bulk queues on Sume
One Sume bulk queue holds 1 to 100 Format runs, so 300 variants means three queues with their own idempotency keys, caps and failure checks. The limits to plan.

A single Sume bulk queue takes 1 to 100 items, so 300 ad variants are three POST .../bulk-runs calls, each with its own Idempotency-Key. Each item is an ordinary Format run with one sandbox, one agent turn and one run receipt. The queue only adds a concurrency window and counts on top.
The limits that shape the plan
| Limit | Value |
|---|---|
| Items per queue | 1 to 100 |
| Concurrency window | 1 to 16 child runs in flight |
| Queue webhook | None; set communication.webhook_url on each item |
| Cancel endpoint for a queue | None; cancel children with POST /v1/format-runs/{run_id}/cancel |
| List endpoint for queues | None; keep the queue ids you get back |
| Keys required | formats:write to create, formats:read to poll; not service-account keys |
Split the work into three queues
Sort variants so each queue is coherent, for example one per hook angle or one per aspect ratio. A failure then costs you a re-run of a slice instead of a re-sort of 300 rows.
Use a stable key per slice, such as a label plus the slice number. A replay of the same key with the same payload returns the existing queue with 202, which makes a retry after a network error safe. The same key with different content is 409 idempotency_conflict. Remember the replay returns the old queue, so mint a new key whenever you really want a fresh batch.
A second reason to split is isolation. If one slice carries a bad assumption, such as a wrong aspect ratio in the instruction, you find out from the first queue's early items and can stop the later slices from ever being created. A single 100 item queue cannot be stopped as a unit, so keep the first slice small when you are testing a new instruction.
Price the worst case before you press go
Every item carries its own generation_spend_cap_usd, and the queue has no total limit. Multiply the number of items by the per-item cap to get the ceiling for each queue, and multiply again by three for the whole job. Three queues of 100 items with a cap of $50 each has a ceiling of $15,000, which is worth knowing before the wallet is checked, not after.
Child runs still pass normal admission, so a wallet that is too small fails an item, not the queue. Such an item becomes failed with run_id: null, and the rest of the queue continues.
Pick a concurrency
The window never goes above the number you send, and workspace generation concurrency still applies to the children. A larger window drains faster but also reaches your spend sooner. Start with a window near the number of runs you are happy to have running while nobody watches, then raise it.
Know when you are done
Poll status_url on each queue, and read counts. A queue is completed when every item is terminal, which does not mean every item succeeded, so branch on counts.failed and counts.canceled. Collect the run_id of each completed item and read the child receipt for the output, because the queue item does not carry it.
Sources
Related posts
More in Formats
- Season outline as nested input: one key, 64-key limit, 2 MiB
Put a whole season outline under one key of a Format run's input. Only top-level keys count toward 64, and the body can be up to 2 MiB, so 40 episodes fit.
- Shorts series sequential playback: fade out or hard cut each episode?
YouTube plays Shorts series episodes in order. Pick a hard cut or a short fade at the end of each one, with the Timeline 1.0 fade limits.
- Bulk queue item error: format_run_canceled vs format_run_failed
In a Sume bulk queue, a failed child reads format_run_failed and a canceled child reads format_run_canceled. The real reason is in the child run receipt.
- TikTok 60-minute uploads vs Sume's 30-minute source limit
Sociality.io reports TikTok uploads up to 60 minutes. Sume trims and Timelines stop at 1800 seconds (30 minutes). What fits, and the cut points for the rest.
Written by Sume