Launching 24 ad variants on Pro: queued jobs, then queue_full

A Pro workspace runs 4 paid generations at once and accepts 24 in total. Plan a TikTok, Reels and Shorts variant launch around those numbers and 429 queue_full.

5 min readSume
All posts

On the Pro plan, Sume processes 4 paid generation jobs at once and holds 20 more in a queue, so 24 variants of a short-form ad are accepted in one launch. The 25th paid submit fails with 429 queue_full. Accepted jobs wait as queued, then move to processing as slots open. Plan the batch around those numbers.

Concurrency is a dispatch limit

The generation admission docs separate four controls that are easy to confuse. Processing concurrency limits jobs in processing. Queue capacity limits jobs that Sume accepted but did not start. Submit rate limits return 429 rate_limited. Balance limits return 402 insufficient_credits before any provider work starts.

That distinction matters for an ad launch. If you submit 24 jobs, you will not get 20 errors. You get 24 durable jobs, 4 running and 20 queued, and you only pay attention to the status.

From the Sume generation admission docs, read 2026-10-05
PlanProcessing concurrencyQueue capacityAccepted capacity
Free156
Pro42024
Startup84048
Scale20100120

What to do with the numbers

Concurrency is plan-only. Prepaid top-ups do not raise it. The default queue capacity is max(3, concurrency_limit x 5). So the way to launch more variants on a smaller plan is to submit in waves, not to add balance.

A good wave size is the accepted capacity of your plan. On Pro that is 24 submits. When some of them finish, submit the next wave. You do not need to wait for all 24.

Submit safely

Give every variant its own Idempotency-Key, built from the campaign, the hook and the aspect ratio. If a submit gets a 429 rate_limited or the connection drops, retry the same body with the same key. You get the same job back, not a second charge.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: fall-launch-hook3-9x16" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "Hook 3: product slides into frame, beat drop, 9:16",
    "duration": 6,
    "aspect_ratio": "9:16",
    "callback_url": "https://example.com/hooks/sume"
  }'

Read status without hammering it

Status and list endpoints have their own rate limits. Treat them as poll backpressure, not as a signal about generation concurrency. A callback_url webhook is the better channel for a launch: the job tells you when it ends, and your poller can sleep.

If a submit fails with queue_full, stop submitting, wait for running jobs to finish, and send that variant again under the same key. Nothing was created, so nothing is duplicated.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume