Testing four new video models at once: how many jobs can Sume hold?

A model bake-off submits many jobs at once. Sume accepts 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale, and queues the rest; past that is 429 queue_full.

4 min readSume
All posts

Sume accepts as many paid generation jobs as your plan's processing concurrency plus its queue capacity: 6 on Free, 24 on Pro, 48 on Startup and 120 on Scale. Past that, new submissions fail with 429 queue_full.

That matters in a launch week, when a bake-off of 5 prompts across 4 models is 20 jobs. These limits come from Sume's generation admission docs, read on 2026-10-07, which also say to prefer the effective field from the API over the static table.

The numbers by plan

Concurrency is a dispatch limit, not a submit limit: a valid job over the cap is accepted as queued and moves to processing when a slot frees. Top-ups do not raise the concurrency limit.

Sume generation limits by plan (read 2026-10-07)
PlanProcessing at onceQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120

A 20-job bake-off on each plan

Twenty jobs fit in the accepted-job capacity of Pro and above. On Free, 6 jobs fit, so the rest must wait for earlier jobs to finish.

Twenty-job bake-off against accepted capacity (read 2026-10-07)
PlanAccepted at onceFits 20 jobs
Free6No: submit in batches
Pro24Yes
Startup48Yes
Scale120Yes

Other limits that bite

  • 429 rate_limited: submit request volume; retry with backoff and an idempotency key.
  • 402 insufficient_credits: Sume reserves the balance at submit, before provider work starts.
  • 429 queue_full: accepted capacity is used up.
  • Poll the status endpoint at a moderate interval; read limits apply there too.

Run a bake-off without waste

Send one Idempotency-Key per prompt and model pair, so a retry after a timeout returns the original job instead of creating a second one. Submit in waves of your accepted capacity, and use the cheapest setting of each row for the first pass.

A wave plan for the Free plan

On Free, 6 jobs are accepted at once, so a 20-job test runs as four waves of 5 or 6. Submit a wave, wait until the jobs complete, then submit the next. The idempotency keys let you resubmit a wave safely if your script restarts partway.

Group each wave by row so a failure points at one model. If a wave returns 429 queue_full, wait for the earlier jobs to finish and try again, rather than retrying in a tight loop.

  • Wave 1: five prompts on the cheapest row.
  • Wave 2 to 4: the other rows, one at a time.
  • Keep the returned job ids in a file for the later download.

Which plan for which test

If your tests are small and occasional, Free works with waves. A regular bake-off of 20 jobs fits Pro without waves, and a team that runs several tests at once will want Startup or Scale. Remember that top-ups add balance, not concurrency, so a plan is the only way to raise the accepted-job number. Read the effective limits from the API for your workspace before you pick, because the static table in the docs is only a guide.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume