Can I submit 100 AI video jobs at once? Queue limits by plan
Accepted capacity is slots plus queue: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Submit 100 at once and 94, 76, 52 or 0 get 429 queue_full.

Sume accepts a paid generation job as long as the workspace has a free processing slot or queue space. Accepted capacity is concurrency plus queue: 6 on Free, 24 on Pro, 48 on Startup and 120 on Scale. If you fire 100 submits at once, the surplus gets 429 queue_full: 94 on Free, 76 on Pro, 52 on Startup and none on Scale.
Where the numbers come from
The docs define the default queue as max(3, concurrency x 5), and accepted capacity as concurrency plus queue. A full processing set is not an error: extra valid jobs wait as queued. Only when the queue is also full does the submit fail.
| Plan | Processing slots | Queue | Accepted capacity | Rejected out of 100 | Wave size hint |
|---|---|---|---|---|---|
| Free | 1 | 5 | 6 | 94 | 4 |
| Pro | 4 | 20 | 24 | 76 | 18 |
| Startup | 8 | 40 | 48 | 52 | 36 |
| Scale | 20 | 100 | 120 | 0 | 90 |
What a rejection costs
Nothing. When admission fails, Sume releases or refunds the reservation, and a rejected job does not capture usage. The error response can include a generation_limits snapshot. The fix is to wait for a job to reach a terminal state, cancel queued jobs you no longer need, and retry with the same idempotency key.
The wave size hint is max(1, floor(queue_capacity_remaining x 0.75)). It is a submission hint, not a concurrency limit, and the docs warn against showing it as one.
A wave planner you can run
This computes how many submit waves 100 jobs need if you stay at the hint size. It makes no network calls; in production, read generation_limits from a real submit response instead of hard-coding the plan table.
import math
def hint(slots: int) -> int:
queue = max(3, slots * 5)
return max(1, ((slots + queue) * 3) // 4)
def main() -> None:
for plan, slots in {"free": 1, "pro": 4, "startup": 8, "scale": 20}.items():
h = hint(slots)
print(plan, "hint", h, "submit waves for 100:", math.ceil(100 / h))
main()The money side
A full queue also means reserved balance: every accepted job reserves its estimated cost at submit time. 24 accepted ten-second 768p H3 Max clips reserve about $24 on Pro; 120 accepted on Scale reserve about $120. A 402 insufficient_credits means the balance could not cover that reserve, and it fails before any provider work starts.
Handling it in practice
A safe loop reads generation_limits from each submit response, stops when queue_capacity_remaining reaches 0, polls the status of open jobs with backoff, and resumes when capacity returns. Use one idempotency key per clip, derived from your own clip id, so a retry after a 429 can never create a second paid job.
On Pro, 100 clips at $1.00 means you can have at most 24 jobs, about $24 of reserve, open at once. After 4 finish, 4 more can be submitted. The total spend is unchanged by the pacing; only the wall clock changes. If you hit queue_full repeatedly, your batch is larger than your plan's pipeline, and the fix is smaller waves, not more retries.
Sources
Related posts
More in Developers
- Read the motion clip length with video inspect before Kling duration
Kling motion control on Sume reserves money from the duration_seconds you declare. Probe the reference clip with video inspect first, so the number is measured.
- Reconcile Sume jobs after a deploy or outage: poll what is open
After downtime, read status for every job your own table still shows as open, honor terminal and result_ready, and never resubmit. Python with sqlite.
- Redact faces and license plates: Pillow first, AI edit only to replace
For redaction use Pillow boxes you control; use an AI mask edit on openai/gpt-image-2.5 only to replace a plate or face, from $0.0094 per image on Sume.
- Redeliver a missed video webhook after a bad deploy: one Sume call
Receiver down when the video finished? POST /v1/jobs/{job_id}/webhook/redeliver re-sends job.completed with a fresh signature. Scope, statuses, pitfalls.
Written by Sume