How many 30-second jobs can you submit at once? Read generation_limits
Submit responses carry generation_limits. Use concurrency_limit minus active and queued jobs as your in-flight budget, and treat wave_size_hint as a hint only.

Your in-flight budget is max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped at queue_capacity_remaining, all read from generation_limits on a submit response. A 30-second video holds a processing slot for minutes, so a batch bigger than that budget just queues behind itself.
Sume accepts valid jobs as queued while queue capacity remains, so full concurrency alone is not an error. The 429 queue_full error appears only when the queue is also full.
Plan limits
Processing concurrency comes from the plan, and prepaid top-ups do not raise it. The default queue capacity is max(3, concurrency_limit x 5). The dashboard Concurrency tab and the generation_limits.concurrency_limit field are the live source of truth, so prefer them to this table.
| Plan | Processing | Queue | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
The wave submitter
This sketch submits a first job, reads its generation_limits, and returns how many more to send. Each new job counts against the budget until you read a fresh snapshot.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def budget(limits: dict) -> int:
free = limits["concurrency_limit"] - limits["active_generation_jobs"] \
- limits["queued_generation_jobs"]
return max(0, min(free, limits["queue_capacity_remaining"]))
def submit(prompt: str, key: str) -> dict:
r = requests.post("https://api.sume.com/v1/video-router/generate",
headers={**H, "Idempotency-Key": key}, timeout=30,
json={"model": "wan-3.0", "prompt": prompt, "duration": 30,
"resolution": "480p", "mode": "async"})
r.raise_for_status()
return r.json()
first = submit("Fog over a harbor at dawn", "harbor-001")
limits = first.get("generation_limits") or first["data"].get("generation_limits")
print("more jobs now:", budget(limits) if limits else "unknown, poll first")When the snapshot is missing
The response includes generation_limits only when Sume can compute it. If it is absent, wait and refresh before you add more work. Never show wave_size_hint as your concurrency, because it counts queue slots too.
Sources
Related posts
More in Developers
- size vs image_size vs aspect_ratio: which field wins on Sume images
On POST /v1/images, image_size beats aspect_ratio, size is a tier word that rejects WxH, and resolution is a tier. Which one to send per model.
- Sora's GET /videos and DELETE /videos/{id}: what Sume offers instead
Sora had list and delete routes. Sume lists your key's own jobs with GET /v1/jobs and documents no public delete, so plan retention in your own storage.
- Sora to Sume in Python: a 10% rollout flag with a spend guard
Move video traffic off a dead Sora call one slice at a time. A stable per-user percentage flag, one Sume call, and a cost guard that stops at your daily cap.
- Speaking rate in words per minute from Sume STT word times (Python)
Compute words per minute for a recording from the words[] start and end times Sume STT returns, plus a per-minute pacing table. Offline Python, no API call.
Written by Sume