Nightly video batch after Sora: 6, 24, 48 or 120 jobs by Sume plan
A cron that submitted 30 Sora renders at once needs a new ceiling. Sume accepts 6 jobs on Free, 24 on Pro, 48 on Startup, 120 on Scale before queue_full.

If your nightly script used to fire a pile of Sora renders in one loop, the number to check before you point it at Sume is accepted job capacity: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale and Enterprise. Submit more than that while none have finished and the extra submits fail with 429 queue_full. Nothing is wrong with the key or the balance.
OpenAI ended the Sora Videos API on 2026-09-24 and its deprecations page names no replacement, per the Pondero report. So the batch has to move to a video API with its own admission rules. Sume's are written down in the generation admission page.
Where the four numbers come from
Sume separates processing concurrency (jobs running now) from queue capacity (accepted jobs waiting). Accepted job capacity is the sum of the two. The default queue capacity is max(3, concurrency_limit x 5).
| Plan | Processing at once | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
What a 30-clip nightly loop does on each plan
Say the loop submits 30 video jobs back to back and the earlier jobs are still running. Free takes 6 and rejects 24. Pro takes 24 and rejects 6. Startup and above take all 30. Those counts assume no job finishes during the loop. Video jobs run for minutes, so for a tight loop that is the realistic case.
A queued job is a normal accepted state, not a failure. The docs say workers move queued jobs into processing as concurrency slots open. Store the job id and poll with backoff.
Read the live number, not the table
Submit responses include a generation_limits object when Sume can compute it. It carries concurrency_limit, queued_jobs_limit, queue_capacity_remaining and wave_size_hint. The docs say to prefer the effective field over the static table, because org workspaces have a floor of 10 concurrent jobs and admin overrides can raise limits. Size each wave from queue_capacity_remaining, then wait for the wave to drain.
On queue_full, the docs tell you to wait for jobs to finish or cancel queued ones, then retry with the same Idempotency-Key. 429 rate_limited is a different thing: request volume on the API itself. Use retry-after when it is present. Do not treat the two the same way in your retry code.
A wave loop outline
The loop below is the shape, with the pieces in the order the docs describe.
- Submit
POST /v1/videoswith a stableIdempotency-Keyper prompt, for example your row id. - Read
generation_limits.queue_capacity_remainingfrom each 202 response and stop submitting at zero. - Poll each job with backoff until
completed,failedorcancelled. - Start the next wave only after capacity returns. Re-send a
queue_fullprompt with the same key.
What this does not change
Prepaid top-ups do not raise processing concurrency; the docs say concurrency is plan-only. A bigger balance does not buy a bigger queue. If the nightly job is now the thing that hits the cap, the choices are a higher plan, smaller waves, or a longer batch window.
Sources
Related posts
More in Developers
- Node 22 script: create a Sume bulk queue and poll it to the end
Dependency-free Node 22 ESM script: POST a bulk queue from items.json, back off the poll, survive 429 and 503, and exit non-zero when any item failed.
- Omni draft grid: four 360p variants, then one final. What it costs
Google's Draft Room idea, run through the Sume API: four 8-second 360p drafts that change one thing each, then a 1080p final. Total $2.70, with a script.
- Pandas DataFrame to a Sume bulk queue in 100-row chunks
Turn a product DataFrame into Sume Format bulk queues: one item per row, 100 rows per queue, a stable key per chunk, and SKU order saved beside each queue id.
- Pin the model id in an ad test: sume/auto follows the catalog
sume/auto is a pure function of the request plus the catalog version, so two ad arms made weeks apart can land on different models. Pin an explicit id in tests.
Written by Sume