Run a 5,000-SKU product video catalog in 50 bulk queues
Split 5,000 SKUs into 50 Formats bulk-run queues of 100, key each shard so a crashed script resumes, and read counts.failed before you call a shard done.

One Formats bulk-run queue holds 1 to 100 items, so a 5,000-SKU catalog is 50 queues of 100. Give each queue a deterministic Idempotency-Key built from a batch label and the shard number, and a crashed script can simply run again: a replayed key returns 202 with the original queue instead of creating a second one.
This post is the shape of that driver. It does not cover prompts, pricing or the Format itself, only how to split, key, poll and retry.
What are the hard numbers?
The limits come from the Sume bulk-runs docs, which defer exact schemas to the live OpenAPI at https://api.sume.com/reference/json.
| Limit | Value | Source |
|---|---|---|
| Items per queue | 1 to 100, in order | Bulk runs, request body |
| Concurrency window | 1 to 16 child runs in flight | Bulk runs, request body |
| Queue webhook | None; webhooks are per item | Bulk runs |
| List or cancel a queue | No public endpoint; cancel a child run | Bulk runs, endpoints |
Queue completed | Every item is terminal, not all succeeded | Bulk runs, queue status |
| Replayed key | 202 with the old queue | Bulk runs |
How do you shard and key 5,000 SKUs?
Slice the SKU list into runs of 100. For shard i, send Idempotency-Key: spring-promo-v1-shard-007 style keys: one fixed label for the batch, plus a zero-padded index. The bulk docs warn that you should mint a fresh key per batch because a spent key replays the old queue. That is the property you want inside a batch and the trap across batches, so change the label when the SKU data or the instruction changes.
Write the shard number, key and returned status_url to a small file or table as you go, so a restart can skip shards you already finished. Keep your own map from shard number to queue id (frq_...). Because there is no list-queues endpoint, that map is how you find your queues again, though replaying the same key also gets you the receipt.
What does the driver look like?
The script below submits shards one at a time, polls each queue until it is completed, and reports counts.failed for the shard. It reads the Format handle and slug from environment variables. Keep the window modest: the docs say workspace generation concurrency still applies to the children, so a window of 16 does not mean 16 children run at once.
import os, time, requests
API = "https://api.sume.com/v1"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
BASE = f"{API}/formats/{os.environ['FORMAT_HANDLE']}/{os.environ['FORMAT_SLUG']}"
LABEL = "spring-promo-v1"
skus = [f"SKU-{n:05d}" for n in range(5000)]
def submit(i, shard):
body = {"concurrency": 8, "items": [
{"instruction": f"Product clip for {s}"} for s in shard]}
r = requests.post(f"{BASE}/bulk-runs", json=body, timeout=60,
headers={**H, "Idempotency-Key": f"{LABEL}-shard-{i:03d}"})
r.raise_for_status()
return r.json()["data"]["status_url"]
for i in range(50):
url = submit(i, skus[i * 100:(i + 1) * 100])
while True:
q = requests.get(url, headers=H, timeout=60).json()["data"]
if q["status"] == "completed":
break
time.sleep(30)
print(i, q["counts"])What happens when a shard fails or the workspace is full?
Two separate things can go wrong, and they look different. The first is a queue that finishes with counts.failed above zero. completed only means every item is terminal, so branch on the counts. The items list gives each row's index, status, run_id and error, so you can build a retry list of just the failed indexes and submit them as a new, differently labelled shard.
The second is admission. The generation-admission docs list 429 queue_full when the workspace has no remaining accepted generation capacity, and 429 rate_limited when request volume exceeds an abuse-protection limit. For queue_full the docs say to wait for jobs to finish or cancel queued jobs and retry with the same idempotency key. For rate_limited they say to back off using retry-after when present. The script above raises on any non-2xx status; wrap submit in a retry for those two codes before a real run. The paced-submit post linked below covers headroom in detail.
How do you size the waves?
Generation submit responses carry a generation_limits object when Sume can compute it, and GET /v1/balance is the other read the docs name for conservative queue decisions. The object reports queue_capacity_remaining and a wave_size_hint, defined in the docs as max(1, floor(queue_capacity_remaining * 0.75)). The docs are explicit that it is a submission-wave hint only. It is not a concurrency limit and must not be used to size in-flight work. The docs do not promise it appears on a bulk-run receipt, so read it from a generation submit response or from the balance read, and slow down while headroom is low rather than racing into queue_full. A queue of 100 items against a small plan will mostly sit in the queued state, which is expected: the window feeds the workspace, and the workspace sets the pace.
- Use 50 shards of 100 items, with one fixed label and a zero-padded index per key.
- Change the label whenever the instruction or SKU data changes, because a spent key replays the old queue.
- Never treat
completedas success; logcounts.failedfor every shard. - Retry only failed indexes, as a new labelled shard.
- Expect no queue-level webhook; poll
status_urland use per-item webhooks if you need pushes. - Check cost on a 100-item shard first, then scale. See the cost-per-SKU post for how that math works.
Sources
Related posts
- Bulk run completed but SKUs failed: retry only the failures
- Pace bulk Sume submits with generation_limits, not wave_size_hint
- Cost per SKU for an AI product video: 500 holiday SKUs priced
- Batch image generation API: how many images can run at once?
- Retry a 503 overload on a paid generation without a double charge
More in Formats
- Revise 20 finished videos in one Sume bulk queue with previous_run_id
Each bulk item can carry previous_run_id, so one queue re-edits twenty finished runs. A preflight for thread_id, the repeated schema, and the cost per item.
- Bulk virtual try-on videos for a catalog: 100 SKUs per queue
Queue up to 100 try-on runs in one POST: concurrency 1 to 16, one garment image per item, a spend cap each, and how to read the failures afterwards.
- Shop Creative Hub limits: 50 per upload, 10k a month, batch math
Creative Hub for GMV Max takes 50 videos per upload and 10k a month. What that means for Sume bulk queues of 100, and where the 200-video link cap bites.
- Shorts ad copy: 40-character headline, 90-character description
Google recommends 40-character headlines and 90-character descriptions for the Shorts ad CTA card. Draft copy in bulk with Sume and check lengths before upload.
Written by Sume