AI video API: rate limit or concurrency, which stops a batch first

Concurrency binds long before request rate on Sume. Free allows 120 writes a minute but 1 running job; Pro 300 and 4. How to size a batch against both.

5 min readSume
All posts

On Sume, the number of video jobs that run at once stops a batch long before the request rate does. A Free key may send 120 writes a minute, but only one video job processes at a time. A Pro key sends 300 writes a minute and processes four. A script that submits 20 clips in one minute never gets near a rate limit; it gets queued, and past the queue it gets 429 queue_full.

The two limits side by side

The authentication docs give the request budget, and the admission docs give the concurrency. Both come from the plan.

Request budget vs generation capacity by plan (Sume docs, read 2026-10-05)
PlanWrites per minuteRunning jobsAccepted jobs (running + queued)
Free12016
Pro300424
Startup600848
Scale120020120

What each failure means

429 rate_limited is request volume. The response carries retry-after, and the error's details.scope names the budget: read or write. 429 queue_full is capacity: the workspace has no accepted generation capacity left until a job finishes or a queued job is canceled. The docs say full concurrency alone is not an error. Valid jobs wait as queued.

Sizing a batch

The admission docs give a rule for new in-flight work: max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. Every submit response carries generation_limits with those counts, and wave_size_hint is a hint for a submission wave, not a concurrency limit. On Pro with an empty workspace, 24 jobs fit and the 25th returns queue_full.

Reads are cheap by design: reads get forty times the write number in their own bucket, so a tight status loop does not use your submit budget. Still, prefer callback_url over polling when you can.

A worked batch

Take 40 clips on a Pro plan. The writes are no issue: 40 submits are under 300 a minute. But the accepted capacity is 24, so the 25th submit returns 429 queue_full. Send 24, wait for jobs to finish, and send more as seats free up. With 4 running at a time, the first 24 take six waves.

The docs also say to send an Idempotency-Key on every unsafe retry, because a retried submit without one can create a second paid job. After a queue_full or a rate_limited response, retry with the same key and back off by retry-after when it is present.

Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.

A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.

If you are new to the API, start with one clip, one model and the lowest resolution, read the full response once, and only then build a loop around it. Most surprises with video jobs come from fields that were defaulted, not from fields that were set.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume