Can I submit 100 AI video jobs at once? Queue limits by plan

Accepted capacity is slots plus queue: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Submit 100 at once and 94, 76, 52 or 0 get 429 queue_full.

4 min readSume
All posts

Sume accepts a paid generation job as long as the workspace has a free processing slot or queue space. Accepted capacity is concurrency plus queue: 6 on Free, 24 on Pro, 48 on Startup and 120 on Scale. If you fire 100 submits at once, the surplus gets 429 queue_full: 94 on Free, 76 on Pro, 52 on Startup and none on Scale.

Where the numbers come from

The docs define the default queue as max(3, concurrency x 5), and accepted capacity as concurrency plus queue. A full processing set is not an error: extra valid jobs wait as queued. Only when the queue is also full does the submit fail.

Accepted job capacity by plan (Sume generation admission docs, read 2026-10-07)
PlanProcessing slotsQueueAccepted capacityRejected out of 100Wave size hint
Free156944
Pro420247618
Startup840485236
Scale20100120090

What a rejection costs

Nothing. When admission fails, Sume releases or refunds the reservation, and a rejected job does not capture usage. The error response can include a generation_limits snapshot. The fix is to wait for a job to reach a terminal state, cancel queued jobs you no longer need, and retry with the same idempotency key.

The wave size hint is max(1, floor(queue_capacity_remaining x 0.75)). It is a submission hint, not a concurrency limit, and the docs warn against showing it as one.

A wave planner you can run

This computes how many submit waves 100 jobs need if you stay at the hint size. It makes no network calls; in production, read generation_limits from a real submit response instead of hard-coding the plan table.

import math


def hint(slots: int) -> int:
    queue = max(3, slots * 5)
    return max(1, ((slots + queue) * 3) // 4)


def main() -> None:
    for plan, slots in {"free": 1, "pro": 4, "startup": 8, "scale": 20}.items():
        h = hint(slots)
        print(plan, "hint", h, "submit waves for 100:", math.ceil(100 / h))


main()

The money side

A full queue also means reserved balance: every accepted job reserves its estimated cost at submit time. 24 accepted ten-second 768p H3 Max clips reserve about $24 on Pro; 120 accepted on Scale reserve about $120. A 402 insufficient_credits means the balance could not cover that reserve, and it fails before any provider work starts.

Handling it in practice

A safe loop reads generation_limits from each submit response, stops when queue_capacity_remaining reaches 0, polls the status of open jobs with backoff, and resumes when capacity returns. Use one idempotency key per clip, derived from your own clip id, so a retry after a 429 can never create a second paid job.

On Pro, 100 clips at $1.00 means you can have at most 24 jobs, about $24 of reserve, open at once. After 4 finish, 4 more can be submitted. The total spend is unchanged by the pacing; only the wall clock changes. If you hit queue_full repeatedly, your batch is larger than your plan's pipeline, and the fix is smaller waves, not more retries.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume