Re-render 300 old Sora prompts: queue capacity by plan
A 300-prompt backlog does not fit in one burst. Use Sume's accepted-job limits per plan, the wave_size_hint, and queue_full handling to pace the re-render.

To re-render a few hundred saved Sora prompts on Sume, pace submissions to your workspace's accepted-job capacity. The generation admission guide lists processing concurrency and queue capacity by plan: for example Pro accepts 24 paid jobs at a time and Scale 120. Past that, submits fail with 429 queue_full. Send a wave, wait for capacity, send the next, and reuse the same Idempotency-Key on any retry.
Know your ceiling
Concurrency is a dispatch limit, not a submit limit: extra valid jobs wait as queued. The queue capacity is the real ceiling for a burst. The numbers below are the documented defaults; the dashboard Concurrency tab and the generation_limits field are the source of truth for your workspace.
| Plan | Processing | Queue | Accepted at once |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
Pace by the snapshot, not a constant
Every submit response includes generation_limits when Sume can compute it. The guide gives a rule for new in-flight work: max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. It also says wave_size_hint is only a hint for submission waves and never a concurrency limit. On Pro with nothing running, the headroom rule gives a budget of 4 in-flight jobs, while the hint (floor of 24 times 0.75) is 18. Pick the conservative number when you want to avoid long queues.
A plan for 300 prompts
Number the prompts, derive a stable Idempotency-Key per prompt, and write the job id next to each row as soon as the submit returns. Submit one wave, poll the batch, and top up as jobs reach a terminal state. On queue_full, stop adding work, cancel queued jobs you no longer need, and retry after capacity opens with the same key.
At 24 accepted jobs, 300 prompts means about 13 waves if each wave fully drains before the next. Real throughput is higher, because you refill as jobs finish.
- Persist the key and job id before the next submit.
- Honor
retry-afterwhen present. - Cancel only works before generation starts; later it returns
409 job_generation_already_started.
Do not skip the money check
A wave also needs balance. Submit fails with 402 insufficient_credits when Sume cannot reserve the estimate, so read GET /v1/balance before each wave. The video generation docs cover the request shape; keep prompts and models in a config file so the job is restartable.
Sources
Related posts
More in Developers
- React Router 7 resource route as a Sume webhook receiver
A React Router 7 route module with only an action is a resource route. Read request.text(), call verifyWebhook from @sume-com/sdk, answer 204, branch on event.
- Debug a migrated video pipeline with Sume job events
When a replaced Sora pipeline stalls, read GET /v1/jobs/{id}/events: created, queued, started, completed or failed, and webhook delivery in one timeline.
- Recraft V4.1 Flash: median 1.3 s, p95 1.8 s. Set timeouts from p95
Recraft quotes a median of about 1.3 seconds and a p95 of 1.8 seconds for V4.1 Flash. How to turn latency claims into timeouts and polling for image APIs.
- Redeliver a missed speech-to-text webhook without rerunning the job
Your receiver was down when a Sume STT job finished. Redeliver the terminal webhook with one call instead of paying to transcribe again.
Written by Sume