Release a 50-SKU holiday batch in two bulk queues: 5 first, 45 after
Submit 5 SKUs as a pilot queue, review the clips, then queue the other 45. Sume bulk queues take 1-100 items at concurrency 1-16. Runnable payload builder.

A Sume bulk queue has no pause button or approval step, so the way to put a person between a pilot and the full batch is to submit two queues. Post 5 of your 50 SKUs as a pilot queue, review the finished clips, and only then post the other 45. Each queue is a server-side list of ordinary Format runs, with concurrency from 1 to 16 and 1 to 100 items, per the bulk runs docs read 2026-10-05. The approval is your policy, not a Sume feature, and it costs nothing extra because each queue is billed run by run.
Why two queues and not one
Once a queue is accepted with 202, the first concurrency items are already in flight, and the API has no public cancel-queue endpoint; cancelling means cancelling each child run. A queue of 50 with concurrency 8 would have eight SKUs spending before you saw a single result. A pilot queue limits what is at risk before review to five runs.
The queue receipt shows completed when every item is terminal, which is not the same as success. Branch on counts.failed before you treat the pilot as clean, as the bulk-run failure guide explains.
The two payloads
The builder below creates both bodies, sets a per-item generation_spend_cap_usd, and checks the documented limits. Use a fresh Idempotency-Key for each queue: replaying a spent key returns 202 with the old queue, which for the second queue would mean it never starts.
import json
SKUS = ["sku-%03d" % n for n in range(1, 51)]
def item(sku):
return {"instruction": "Product clip for " + sku,
"input": {"sku": sku},
"generation_spend_cap_usd": 8}
def queue(skus, concurrency):
assert 1 <= concurrency <= 16 and 1 <= len(skus) <= 100
return {"concurrency": concurrency, "items": [item(s) for s in skus]}
pilot = queue(SKUS[:5], concurrency=5)
rest = queue(SKUS[5:], concurrency=8)
print(len(pilot["items"]), len(rest["items"]))
print(json.dumps(pilot)[:100])
What to review between queues
Poll GET /v1/format-run-queues/{id} until the pilot is terminal, then read each child at GET /v1/format-runs/{run_id}. A per-item communication.webhook_url gives you a signed delivery for each SKU as it finishes, as in the webhook review queue.
| Step | Queue | Items | Concurrency | Gate |
|---|---|---|---|---|
| 1 | Pilot | 5 | 5 | Reviewer approves clips and counts.failed is 0 |
| 2 | Remainder | 45 | 8 | Poll counts, requeue failures |
| 3 | Retry queue | Failed SKUs only | Up to 16 | New Idempotency-Key |
Choose the pilot SKUs to be the hard ones, such as reflective or transparent products, not the easy ones, so a pass means something. Workspace generation concurrency still applies to the children, so a larger concurrency value does not run faster than your plan allows. For batches larger than 100, see the 250-SKU split into three queues.
Sources
Related posts
More in Formats
- Format run, schedule or Agent Completion for repeatable holiday ads
Use a Format when only inputs change, Scheduled when the clock starts the work, Agent Completions when the task changes. Set a spend cap on each.
- Retry only the missing ad variants with previous_run_id on a Format
When a variants run returns some clips as null, continue it with previous_run_id and name the missing slots, so you do not pay for the finished variants again.
- Retry two scenes in one Format run: scene_ids with previous_run_id
Continue the thread with previous_run_id and put scene_ids in input to redo two clips in one run. Voice and other clips stay; send a cap.
- Revoke a Format grant: in-flight runs finish, new calls fail closed
DELETE .../grants/{workspace} revokes a pending or accepted grant. The grantee's running runs finish, and every new list, read or invoke fails closed at once.
Written by Sume