Nightly video batch after Sora: 6, 24, 48 or 120 jobs by Sume plan

A cron that submitted 30 Sora renders at once needs a new ceiling. Sume accepts 6 jobs on Free, 24 on Pro, 48 on Startup, 120 on Scale before queue_full.

5 min readSume
All posts

If your nightly script used to fire a pile of Sora renders in one loop, the number to check before you point it at Sume is accepted job capacity: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale and Enterprise. Submit more than that while none have finished and the extra submits fail with 429 queue_full. Nothing is wrong with the key or the balance.

OpenAI ended the Sora Videos API on 2026-09-24 and its deprecations page names no replacement, per the Pondero report. So the batch has to move to a video API with its own admission rules. Sume's are written down in the generation admission page.

Where the four numbers come from

Sume separates processing concurrency (jobs running now) from queue capacity (accepted jobs waiting). Accepted job capacity is the sum of the two. The default queue capacity is max(3, concurrency_limit x 5).

Sume generation limits by plan, default queue (read 2026-10-07)
PlanProcessing at onceQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

What a 30-clip nightly loop does on each plan

Say the loop submits 30 video jobs back to back and the earlier jobs are still running. Free takes 6 and rejects 24. Pro takes 24 and rejects 6. Startup and above take all 30. Those counts assume no job finishes during the loop. Video jobs run for minutes, so for a tight loop that is the realistic case.

A queued job is a normal accepted state, not a failure. The docs say workers move queued jobs into processing as concurrency slots open. Store the job id and poll with backoff.

Read the live number, not the table

Submit responses include a generation_limits object when Sume can compute it. It carries concurrency_limit, queued_jobs_limit, queue_capacity_remaining and wave_size_hint. The docs say to prefer the effective field over the static table, because org workspaces have a floor of 10 concurrent jobs and admin overrides can raise limits. Size each wave from queue_capacity_remaining, then wait for the wave to drain.

On queue_full, the docs tell you to wait for jobs to finish or cancel queued ones, then retry with the same Idempotency-Key. 429 rate_limited is a different thing: request volume on the API itself. Use retry-after when it is present. Do not treat the two the same way in your retry code.

A wave loop outline

The loop below is the shape, with the pieces in the order the docs describe.

  • Submit POST /v1/videos with a stable Idempotency-Key per prompt, for example your row id.
  • Read generation_limits.queue_capacity_remaining from each 202 response and stop submitting at zero.
  • Poll each job with backoff until completed, failed or cancelled.
  • Start the next wave only after capacity returns. Re-send a queue_full prompt with the same key.

What this does not change

Prepaid top-ups do not raise processing concurrency; the docs say concurrency is plan-only. A bigger balance does not buy a bigger queue. If the nightly job is now the thing that hits the cap, the choices are a higher plan, smaller waves, or a longer batch window.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume