8 avatar videos on a Free Sume workspace: 6 accepted, then 429
A Free Sume workspace runs 1 job, queues 5 and accepts 6. The 7th and 8th avatar video submits get 429 queue_full. A retry loop that waits and reuses the key.
On a Free Sume workspace, a loop that submits 8 avatar videos in a row gets 6 accepted jobs and a 429 queue_full on the seventh. The generation-admission docs give Free a processing concurrency of 1 and a queue capacity of 5, so accepted job capacity is 6 (read 2026-10-09). One job renders while five wait, and the rest are refused until a slot or queue place opens.
What each response means
Concurrency alone is never an error. The docs say Sume accepts valid jobs as queued while queue capacity remains. A 429 appears when both the processing slots and the queue are full. 429 rate_limited is a different failure: too many requests in a window. The two need different handling.
| Submit | State after submit | Response |
|---|---|---|
| 1 | processing | accepted |
| 2 to 6 | queued | accepted |
| 7 | queue full | 429 queue_full |
| 8 | queue full | 429 queue_full |
A retry loop that does not double-charge
The docs say to retry a queue_full with the same idempotency key after waiting for jobs to finish or canceling queued ones, and never to retry an unsafe submit without a key. The loop below sets one key per clip and honors retry-after when the header is present. Note that queue_full may not carry that header in every case, so it falls back to 30 seconds. At 8 clips that never succeeds on a Free plan until earlier jobs finish, which is the point: the loop waits instead of dropping work.
import os, time, requests
URL = "https://api.sume.com/v1/avatar-1.0/talking-video"
HEAD = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json"}
def submit(i, script):
body = {"avatar_handle": "studio_presenter", "script": script,
"quality": "standard"}
h = dict(HEAD, **{"Idempotency-Key": f"welcome-{i}"})
while True:
r = requests.post(URL, json=body, headers=h, timeout=30)
if r.status_code == 429:
time.sleep(int(r.headers.get("retry-after", "30")))
continue
r.raise_for_status()
return r.json()
scripts = [f"Welcome, customer {n}." for n in range(8)]
jobs = [submit(n, s) for n, s in enumerate(scripts)]
print(len(jobs), "submitted")Choosing a better shape
For a large batch on a Free workspace, submit in waves of 6, poll GET /v1/jobs/:id/status until they complete, then submit the next. Eight clips at standard 20 seconds is 8 x $3.68 = $29.44, and the first six alone are $22.08. If you routinely need more than six at a time, a higher plan raises both numbers: Pro accepts 24 and Scale 120.
Remember that top-ups do not raise concurrency. Only the plan, or an admin override, changes it, and the response's generation_limits block shows the effective limit for your workspace.
Why not cancel and resubmit
It is tempting to cancel queued jobs to make room. The docs allow canceling queued work, and a job canceled before generation starts is refunded. But cancel only what you truly do not want. A canceled job loses its place, and resubmitting with a new key creates a new job. Reusing the same key for an identical retry is the safe path; using it for a changed payload is a conflict (409 idempotency_conflict).
Also keep reads gentle. Status polling has its own rate limits, which the docs describe as poll backpressure, not generation concurrency. Poll each job every few seconds with backoff instead of in a tight loop.
Sources
Related posts
More in Developers
- Eight Wan 3.0 clips on Free: six queue, two get queue_full
Free plan accepted capacity is 6 paid jobs. Submit eight 10-second Wan 3.0 720p jobs and two return 429 queue_full; the six accepted hold $7.50 of the $10.00.
- A 7-second clip on six Sume video models: $0.525 to $4.0446
Per-second list prices times 7 s for six Sume video models, with each model's accepted duration range and a Python check that rejects out-of-range lengths.
- A fresh UUID per retry is not an idempotency key (Node, Sume)
Create the Idempotency-Key once outside the retry loop, retry only 429, 502, 503 and 504, and stop on 4xx. A Node submit helper for Sume /v1/videos.
- Per-request timeout on Sume video polls: 15 s abort, then poll again
Give every poll GET its own 15 s AbortSignal.timeout so one hung read does not freeze a 30 s Wan or Seedance job loop. Reads per job per hour at 5, 10, 30 s.
Written by Sume