One Format run with five variants or a five-item bulk queue?
One run gives you five variants under one cap and one receipt. A bulk queue gives five receipts, a concurrency window up to 16 and per-item retry.

Use one Format run with a variants[] schema when the five variants share a brief and you want them as a set. Use a bulk queue of five items when each variant is its own job, needs its own cap or webhook, or must retry alone. A bulk request is a server-side queue of ordinary runs, so the difference is how many agent turns and receipts you get.
One run, five variants
One run is one sandbox and one agent turn. The agent sees the whole set at once, which helps with contrast between hooks, and one spend cap bounds the lot. The cost is a shared failure domain. If the run stops, every unfinished variant stops with it, and the only record is the partial output.
A queue of five items
A queue takes concurrency from 1 to 16 and 1 to 100 items, each one a full run body. The server keeps concurrency children in flight and starts the next as a slot frees. Each child binds its own output_schema, spend cap and webhook. The queue itself has no webhook, so you poll GET /v1/format-run-queues/{id} or take each child's terminal delivery.
| Question | One run, variants[] | Bulk queue, 5 items |
|---|---|---|
| Agent turns | 1 | 5 |
| Receipts | 1 | 5 child receipts plus the queue |
| Spend cap | One per run | One per item |
| Webhook | One terminal delivery | One per child; none on the queue |
| Retry one variant | Continue with previous_run_id | Cancel or re-send that item |
| Parallelism | Inside the agent turn | Window of 1 to 16 |
| Failure | Whole run fails together | Item fails, queue goes on |
Reading a finished queue
Queue status completed means every item is terminal, not that every item succeeded. Branch on counts.failed and counts.canceled, and read a failed child's receipt for the reason. Use a fresh Idempotency-Key for each batch, because replaying a spent key returns the old queue with 202.
The queue request
The queue body for five hook variants is small. Each item carries its own hook in input, and the hook is data, so keep it out of the instruction. The script prints the queue id and the count of running children from the 202 receipt.
import json, os, urllib.request
HOOKS = ["Stop scrolling", "Three things I wish I knew", "Before and after",
"Unboxing in ten seconds", "The one setting that matters"]
def main():
items = [{"instruction": "Make one vertical ad using the hook in input.",
"input": {"hook": h}, "generation_spend_cap_usd": 5} for h in HOOKS]
body = {"concurrency": 3, "items": items}
req = urllib.request.Request(
"https://api.sume.com/v1/formats/sume/sume-video-hook/bulk-runs",
data=json.dumps(body).encode(), method="POST",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": "hooks-batch-2026-10-05"})
with urllib.request.urlopen(req) as r:
q = json.load(r)["data"]
print(q["id"], q.get("counts"))
main()A rule of thumb
A simple rule works. If the variants are one creative decision, run one. If they are five independent deliverables going to five places, queue five. When unsure, start with the queue, because per-item caps and retries are easier to reason about than a shared turn.
Sources
Related posts
More in Formats
- Can a partner bulk-run your shared Format? Queue and spend are theirs
A grantee can POST .../bulk-runs at your handle and slug with its own team key. The queue, child runs and spend are theirs, and you cannot poll their queue.
- Price-drop sale videos for 200 SKUs: two bulk queues, one key each
A bulk queue holds at most 100 items. For 200 marked-down SKUs, send two queues, give each its own Idempotency-Key, and track both queue ids yourself.
- Release a 50-SKU holiday batch in two bulk queues: 5 first, 45 after
Submit 5 SKUs as a pilot queue, review the clips, then queue the other 45. Sume bulk queues take 1-100 items at concurrency 1-16. Runnable payload builder.
- Format run, schedule or Agent Completion for repeatable holiday ads
Use a Format when only inputs change, Scheduled when the clock starts the work, Agent Completions when the task changes. Set a spend cap on each.
Written by Sume