Testing four new video models at once: how many jobs can Sume hold?
A model bake-off submits many jobs at once. Sume accepts 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale, and queues the rest; past that is 429 queue_full.

Sume accepts as many paid generation jobs as your plan's processing concurrency plus its queue capacity: 6 on Free, 24 on Pro, 48 on Startup and 120 on Scale. Past that, new submissions fail with 429 queue_full.
That matters in a launch week, when a bake-off of 5 prompts across 4 models is 20 jobs. These limits come from Sume's generation admission docs, read on 2026-10-07, which also say to prefer the effective field from the API over the static table.
The numbers by plan
Concurrency is a dispatch limit, not a submit limit: a valid job over the cap is accepted as queued and moves to processing when a slot frees. Top-ups do not raise the concurrency limit.
| Plan | Processing at once | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
A 20-job bake-off on each plan
Twenty jobs fit in the accepted-job capacity of Pro and above. On Free, 6 jobs fit, so the rest must wait for earlier jobs to finish.
| Plan | Accepted at once | Fits 20 jobs |
|---|---|---|
| Free | 6 | No: submit in batches |
| Pro | 24 | Yes |
| Startup | 48 | Yes |
| Scale | 120 | Yes |
Other limits that bite
429 rate_limited: submit request volume; retry with backoff and an idempotency key.402 insufficient_credits: Sume reserves the balance at submit, before provider work starts.429 queue_full: accepted capacity is used up.- Poll the status endpoint at a moderate interval; read limits apply there too.
Run a bake-off without waste
Send one Idempotency-Key per prompt and model pair, so a retry after a timeout returns the original job instead of creating a second one. Submit in waves of your accepted capacity, and use the cheapest setting of each row for the first pass.
A wave plan for the Free plan
On Free, 6 jobs are accepted at once, so a 20-job test runs as four waves of 5 or 6. Submit a wave, wait until the jobs complete, then submit the next. The idempotency keys let you resubmit a wave safely if your script restarts partway.
Group each wave by row so a failure points at one model. If a wave returns 429 queue_full, wait for the earlier jobs to finish and try again, rather than retrying in a tight loop.
- Wave 1: five prompts on the cheapest row.
- Wave 2 to 4: the other rows, one at a time.
- Keep the returned job ids in a file for the later download.
Which plan for which test
If your tests are small and occasional, Free works with waves. A regular bake-off of 20 jobs fits Pro without waves, and a team that runs several tests at once will want Startup or Scale. Remember that top-ups add balance, not concurrency, so a plan is the only way to raise the accepted-job number. Read the effective limits from the API for your workspace before you pick, because the static table in the docs is only a guide.
Sources
Related posts
More in Developers
- TTS output formats: Gemini 24 kHz WAV vs Sume 44.1 kHz MP3 default
Gemini 3.8 TTS returns 24 kHz mono 16-bit WAV, or headerless L16 when streaming. Sume TTS defaults to 44.1 kHz 128 kbps MP3 and offers WAV and raw options.
- Sume TTS tts_duration_exceeded: splitting a script over 1,200 seconds
Sume TTS fails audio longer than 1,200 seconds with tts_duration_exceeded, and caps a request at 20,000 characters. Split at paragraph breaks and rejoin.
- Sume TTS 409 tts_voice_language_mismatch: confirm and retry safely
A Sume TTS 409 means the voice's language differs from your request. Nothing was charged. Ask the user, then retry with confirm_language_mismatch true.
- $2.00 balance: one 10-second Omni 1080p job, then a 402
With $2.00 in the wallet, one 10-second Omni 1080p job reserves $1.875. A second identical submit gets 402, not 429: balance and queue are separate.
Written by Sume