Python fallback chain for Sume video models: 404 and 503 only
After the Sora API shutdown, a model chain must not retry everything. This Python function moves on at 404 and 503 and stops on 429, 402, 400, 409.

A video model fallback chain should move to the next model on only two answers, 404 (the model id is not found) and 503 (provider capacity), and stop on everything else. A 429 means wait, a 402 means add funds, a 400 means your request is wrong, and a 409 means your key was reused. A different model fixes none of those. The Python below implements that rule and gives each model its own idempotency key.
The reason to build this now is the OpenAI deprecations page. It lists the Videos API and six Sora 2 ids as removed on 2026-09-24 and names no replacement.
What the vendor page says
Because the vendor lists no replacement, the pipeline has to pick one. A chain with an ordered model list is a way to make that choice a config value instead of a code change.
| Item | Value |
|---|---|
| Announced | 2026-03-24 |
| Shutdown date | 2026-09-24 |
| Listed ids | sora-2, sora-2-pro, sora-2-2025-10-06, sora-2-2025-12-08, sora-2-pro-2025-10-06, plus the Videos API |
| Recommended replacement | None listed |
The chain
The key for each model is a hash of the job key, the model, and the prompt, cut to 32 characters. The same model and prompt always give the same key, so a retry of the whole function after a crash replays the accepted job instead of billing a second one. A different model gets a different key, which avoids a 409 idempotency_conflict caused by a changed payload.
import hashlib, json, os, urllib.error, urllib.request
BASE = os.environ.get("SUME_BASE", "https://api.sume.com")
def post(model, prompt, key):
req = urllib.request.Request(
BASE + "/v1/videos", method="POST",
data=json.dumps({"model": model, "prompt": prompt}).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json", "Idempotency-Key": key})
try:
with urllib.request.urlopen(req, timeout=35) as r:
return r.status, json.load(r)
except urllib.error.HTTPError as e:
return e.code, json.load(e)
def submit_with_fallback(models, prompt, job_key):
for model in models:
key = hashlib.sha256(f"{job_key}|{model}|{prompt}".encode()).hexdigest()[:32]
status, body = post(model, prompt, key)
if status == 202:
return model, body
if status not in (404, 503): # 429, 402, 400, 409: switching models does not help
raise RuntimeError(f"{status} {body.get('error', {}).get('code')} on {model}")
print(f"{model}: {status}, trying next")
raise RuntimeError("every model in the chain refused")
print(submit_with_fallback(["old-model", "ok-model"], "A mug on a desk", "order-8823"))What the stub run showed
On a local stub, old-model answered 404 and ok-model answered 202. The function printed old-model: 404, trying next and returned the second model with its job envelope. A 429 on the first model would have raised at once, which is the point.
A 503 may clear in a few seconds. If you prefer, wait once and retry the same model before you move on; the docs say to retry later with the same key for provider_capacity_exceeded.
Before you trust the chain
- Check each model's duration, resolution, and ratio limits against the catalog. A fallback that does not accept the same clip will fail with a 400.
- Log which model served each job. Output differs between models, and your team should know which one it got.
- Test the chain with a stub, not with paid jobs.
Putting the chain in config
Keep the model list outside the code, for example as a comma-separated environment variable, and split it at start. The first entry is the model you want; the rest are the ones you accept. A change of vendor, or a model that leaves the catalog, then needs a config edit and a restart, and not a deploy.
Order matters for cost as well as quality. Sume bills the provider list price times 1.25 for each model, so a fallback can cost more or less than the first choice. Read the price of each entry before you rely on it, and set a spend limit on the batch.
sume/auto is another option: you send it as the model, and Sume selects the family. It does not tell you which family ran, so use a pinned chain when you need to know.
Sources
Related posts
More in Developers
- Python: find Sume image models that list a ratio like 8:1 or 4:5
A 15-line Python script reads GET /v1/images/models and prints every Sume image model that lists a given aspect ratio, so you stop guessing before a 400.
- Python match on a Sume /v1/videos poll: five statuses, one handler
Python 3.10 structural matching on the poll dict: wait on pending and in_progress, return on completed, raise on failed or cancelled. Runs under asyncio.run.
- Parse Retry-After as seconds or HTTP-date before retrying a Sume 429
A 21-line Python helper that reads retry-after as integer seconds or an HTTP-date, caps the wait, and falls back to exponential delay when the header is absent.
- Save a Sume job's result artifacts in Python by content type
Fetch GET /v1/jobs/:id/result and save each artifact with an extension from content_type, not the URL. Standard library only, with a text-result guard.
Written by Sume