Run a 5,000-SKU product video catalog in 50 bulk queues

Split 5,000 SKUs into 50 Formats bulk-run queues of 100, key each shard so a crashed script resumes, and read counts.failed before you call a shard done.

5 min readSume
All posts

One Formats bulk-run queue holds 1 to 100 items, so a 5,000-SKU catalog is 50 queues of 100. Give each queue a deterministic Idempotency-Key built from a batch label and the shard number, and a crashed script can simply run again: a replayed key returns 202 with the original queue instead of creating a second one.

This post is the shape of that driver. It does not cover prompts, pricing or the Format itself, only how to split, key, poll and retry.

What are the hard numbers?

The limits come from the Sume bulk-runs docs, which defer exact schemas to the live OpenAPI at https://api.sume.com/reference/json.

Bulk-run limits that shape a 5,000-SKU plan, read 2026-10-03
LimitValueSource
Items per queue1 to 100, in orderBulk runs, request body
Concurrency window1 to 16 child runs in flightBulk runs, request body
Queue webhookNone; webhooks are per itemBulk runs
List or cancel a queueNo public endpoint; cancel a child runBulk runs, endpoints
Queue completedEvery item is terminal, not all succeededBulk runs, queue status
Replayed key202 with the old queueBulk runs

How do you shard and key 5,000 SKUs?

Slice the SKU list into runs of 100. For shard i, send Idempotency-Key: spring-promo-v1-shard-007 style keys: one fixed label for the batch, plus a zero-padded index. The bulk docs warn that you should mint a fresh key per batch because a spent key replays the old queue. That is the property you want inside a batch and the trap across batches, so change the label when the SKU data or the instruction changes.

Write the shard number, key and returned status_url to a small file or table as you go, so a restart can skip shards you already finished. Keep your own map from shard number to queue id (frq_...). Because there is no list-queues endpoint, that map is how you find your queues again, though replaying the same key also gets you the receipt.

What does the driver look like?

The script below submits shards one at a time, polls each queue until it is completed, and reports counts.failed for the shard. It reads the Format handle and slug from environment variables. Keep the window modest: the docs say workspace generation concurrency still applies to the children, so a window of 16 does not mean 16 children run at once.

import os, time, requests

API = "https://api.sume.com/v1"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
BASE = f"{API}/formats/{os.environ['FORMAT_HANDLE']}/{os.environ['FORMAT_SLUG']}"
LABEL = "spring-promo-v1"
skus = [f"SKU-{n:05d}" for n in range(5000)]

def submit(i, shard):
    body = {"concurrency": 8, "items": [
        {"instruction": f"Product clip for {s}"} for s in shard]}
    r = requests.post(f"{BASE}/bulk-runs", json=body, timeout=60,
        headers={**H, "Idempotency-Key": f"{LABEL}-shard-{i:03d}"})
    r.raise_for_status()
    return r.json()["data"]["status_url"]

for i in range(50):
    url = submit(i, skus[i * 100:(i + 1) * 100])
    while True:
        q = requests.get(url, headers=H, timeout=60).json()["data"]
        if q["status"] == "completed":
            break
        time.sleep(30)
    print(i, q["counts"])

What happens when a shard fails or the workspace is full?

Two separate things can go wrong, and they look different. The first is a queue that finishes with counts.failed above zero. completed only means every item is terminal, so branch on the counts. The items list gives each row's index, status, run_id and error, so you can build a retry list of just the failed indexes and submit them as a new, differently labelled shard.

The second is admission. The generation-admission docs list 429 queue_full when the workspace has no remaining accepted generation capacity, and 429 rate_limited when request volume exceeds an abuse-protection limit. For queue_full the docs say to wait for jobs to finish or cancel queued jobs and retry with the same idempotency key. For rate_limited they say to back off using retry-after when present. The script above raises on any non-2xx status; wrap submit in a retry for those two codes before a real run. The paced-submit post linked below covers headroom in detail.

How do you size the waves?

Generation submit responses carry a generation_limits object when Sume can compute it, and GET /v1/balance is the other read the docs name for conservative queue decisions. The object reports queue_capacity_remaining and a wave_size_hint, defined in the docs as max(1, floor(queue_capacity_remaining * 0.75)). The docs are explicit that it is a submission-wave hint only. It is not a concurrency limit and must not be used to size in-flight work. The docs do not promise it appears on a bulk-run receipt, so read it from a generation submit response or from the balance read, and slow down while headroom is low rather than racing into queue_full. A queue of 100 items against a small plan will mostly sit in the queued state, which is expected: the window feeds the workspace, and the workspace sets the pace.

  • Use 50 shards of 100 items, with one fixed label and a zero-padded index per key.
  • Change the label whenever the instruction or SKU data changes, because a spent key replays the old queue.
  • Never treat completed as success; log counts.failed for every shard.
  • Retry only failed indexes, as a new labelled shard.
  • Expect no queue-level webhook; poll status_url and use per-item webhooks if you need pushes.
  • Check cost on a 100-item shard first, then scale. See the cost-per-SKU post for how that math works.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume