Startup plan accepts 48 jobs, not 50: send 50 AI video clips in waves
Sume admits concurrency plus queue jobs per plan: 6, 24, 48 or 120. Fifty 30-second clips fit only on Scale at once. See waves per plan and how to retry.

Only a Scale workspace can hold 50 paid video jobs at once: accepted capacity is 6 on Free, 24 on Pro, 48 on Startup and 120 on Scale. On a smaller plan the 49th or 25th submit returns 429 queue_full, so send the clips in waves and retry with the same Idempotency-Key when jobs finish.
For 30-second Seedance 2.5 and Wan 3.0 renders the point is more practical than it sounds. Those jobs hold their processing slot for minutes, so the queue stays full for a while and you should plan for it.
Concurrency is not the submit limit
Sume accepts valid jobs as queued while queue capacity remains, and workers move them to processing under the plan's concurrency limit. The default queue capacity is max(3, concurrency_limit x 5), and accepted capacity is concurrency plus queue. Full concurrency alone is not an error; it becomes 429 queue_full only when the queue is also full (Generation admission).
| Plan | Processing | Queue | Accepted | wave_size_hint | Waves for 50 |
|---|---|---|---|---|---|
| Free | 1 | 5 | 6 | 4 | 13 |
| Pro | 4 | 20 | 24 | 18 | 3 |
| Startup | 8 | 40 | 48 | 36 | 2 |
| Scale | 20 | 100 | 120 | 90 | 1 |
How the wave sizes are computed
wave_size_hint is max(1, floor(queue_capacity_remaining x 0.75)) on an empty workspace, so Free is floor(6 x 0.75) = 4, Pro is 18, Startup is 36, and Scale is 90. Treat it only as a hint for how many to submit together, never as the concurrency limit. The effective concurrency_limit in the generation_limits object is the authority, because admin overrides can differ from the table.
Waves for 50 are ceil(50 / hint). They count submissions, not elapsed time. Sume does not publish a per-job queue position or an ETA, so do not promise one.
What to do on queue_full
- Do not add more work for that workspace.
- Poll your current jobs until at least one is terminal, with backoff.
- Cancel queued jobs you no longer need. Cancel works only before generation starts.
- Retry the rejected submit with the same Idempotency-Key, using
retry-afterwhen present.
A bounded submitter
This loop pauses on queue_full, then continues with the same key. It stops cleanly on a 402 instead of looping.
import os, time, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def submit(key, body):
while True:
r = requests.post("https://api.sume.com/v1/videos", json=body,
headers={**H, "Idempotency-Key": key}, timeout=30)
if r.status_code == 402:
raise SystemExit("balance stop: " + r.text)
if r.status_code == 429:
time.sleep(float(r.headers.get("retry-after", 30)))
continue
r.raise_for_status()
return r.json()["id"]
ids = [submit(f"promo-v2-{i:02d}", {"model": "wan-3.0", "duration": 30,
"resolution": "480p", "prompt": f"Scene {i}"}) for i in range(50)]
print(len(ids), "accepted")Money check
Fifty 30-second Wan 3.0 clips at 480p are $93.75 on Sume, so the balance must cover about that before the first submit. The reservation is taken at accept time, which is why a half-submitted batch can stop on 402 instead of queue_full.
Checking the number you were given
Do not hard-code 48. Read wave_size_hint from the admission information Sume returns and size each wave from it. The hint is the floor of the remaining queue capacity times 0.75, so it changes as your queue drains and as your plan changes.
After each wave, wait for most jobs to finish before you submit the next one. A 429 queue_full is a signal to wait, not to retry in a tight loop, and your submit helper should treat it that way.
Sources
Related posts
More in Developers
- Stitch 10-second AI clips into a 3-minute Short with Timeline 1.0
YouTube's AI Shorts clips top out at 10 seconds, but a Short can run 3 minutes. Join clips and a voice track in Timeline 1.0 within its 200 slots and 1,800 s.
- Stitch AI clips into one MP4 with Sume Timeline in Python
POST /v1/timeline-1.0/render with audio mode silence joins clips into one MP4 for $0.10 a minute. A short Python script, with the plan call first.
- Stitch three STT chunks: add each range start to word times
Sume STT times count from each chunk's own start, so add the chunk's range start to every word and sentence. A short Python function that does it.
- Stop an Omni draft batch at $5: sum usage.cost from each poll
A Python loop that submits 360p Omni drafts one at a time, adds the Sume usage.cost of each finished job, and stops before the next one would cross a $5 cap.
Written by Sume