POST /v1/videos status codes: which of 10 are safe to retry
OpenAPI lists 202 plus ten error codes for POST /v1/videos. Which to retry with the same key, which to fix or stop on, a Python classifier and a backoff plan.

POST /v1/videos documents 202 plus ten error codes: 400, 401, 402, 404, 409, 413, 429, 500, 502 and 503. Retry only 429, 500, 502 and 503, always with the same Idempotency-Key; fix and resend for 400, 404 and 413; stop for 401; top up then resend for 402; and adopt the existing job on 409.
Retry map
The 502 body reads that Sume could not start the generation job. A 5xx never proves the job was refused, which is why the same key matters: with a key, the retry adopts whatever the first call created instead of paying for it twice.
| Status | Meaning | Retry? | Action |
|---|---|---|---|
| 202 | Accepted, job created | No | Store id and polling_url |
| 400 | Validation, unsupported_parameter | No | Fix the body |
| 401 | Key missing or invalid | No | Fix the credential; send only one |
| 402 | insufficient_credits, no job started | After top-up | Add balance, resend with the same key |
| 404 | model_not_found | No | Pick an id from GET /v1/videos/models |
| 409 | idempotency_conflict | No | Adopt the job in error.details |
| 413 | Body too large | No | Shrink inputs |
| 429 | queue_full or rate_limited | Yes | Wait for retry-after |
| 500 | Server error | Yes | Same key |
| 502 | Could not start the generation job | Yes | Same key |
| 503 | provider_capacity_exceeded or overload | Yes | Same key |
A classifier
Keep the decision in one function so every call site agrees. The wait is the retry-after header when present, else exponential: 2, 4, 8, 16 s, which adds up to 30 s over four retries.
RETRY = {429, 500, 502, 503}
def next_step(status, headers, attempt):
if status == 202:
return "done", 0
if status in RETRY and attempt < 4:
wait = headers.get("retry-after")
return "retry", int(wait) if wait else 2 ** (attempt + 1)
if status == 409:
return "adopt", 0 # read error.details.job_id
if status == 402:
return "top_up", 0
return "stop", 0Gotchas
A 429 names the budget in error.details.scope (read or write); queue_full is about generation capacity, not request rate. Cancel is a different story: it works only before generation starts and returns 409 job_generation_already_started afterwards.
Retrying 500, 502 and 503 without a key is the expensive mistake: at 30 s of Wan 3.0 at 720p, three unkeyed attempts can reserve 3 x $3.75 = $11.25.
Sources
Related posts
More in Developers
- Preflight a mixed 30 s batch: 2 Seedance and 4 Wan needs $49.67
Sum the reserve for two Seedance 2.5 720p clips and four Wan 3.0 720p clips (49,668,000 micros), compare with GET /v1/balance in integers, then submit.
- Fresh Idempotency-Key per proxy call: why a Sume retry bills twice
If your server proxy mints a new Idempotency-Key on every request, a browser retry becomes a second paid Sume job. Forward the client's key instead; TypeScript.
- Python asyncio loop for a Sume job: next_poll_after_seconds
A runnable httpx and asyncio loop for GET /v1/jobs/{id}/status that honors next_poll_after_seconds, backs off otherwise, and leaves the job running on timeout.
- Python cost cap for a mixed media job: round up, then refuse
A 25-line Python estimator with Decimal prices that rounds the total up to the cent, as Sume's reservation does, and exits before submit if it is over your cap.
Written by Sume