Video model fallback chain in Python with one key per model
A Python chain that tries Gemini Omni, Wan 3.0 and Seedance 2 on Sume in order, stops on 401 or 402, and uses a separate idempotency key for each model.

Write the fallback as a loop over catalog model ids, send each attempt with its own Idempotency-Key, and stop the whole chain on 401 or 402, because every later model would fail the same way. The Python below tries gemini-omni-flash-1.1, then wan-3.0, then seedance-2, and returns the first completed file.
This matters now because OpenAI's deprecations page lists the Sora video models and the Videos API as removed on 2026-09-24 with no replacement named. A pipeline that had one video vendor has one video vendor fewer, and the lesson is to never hard-wire a single model id again.
Why the key is per model
Sume's idempotency rule is that a replay with the same key and the same body returns the original job, while the same key with a different body returns 409 conflict. A different model is a different body. If you reuse one key across the chain, the second attempt collides with the first and you get a 409 instead of a render.
Building the key as the row id plus the model id fixes both sides. A crash and restart replays each attempt to its original job, so you never pay twice for the same row and model. A fallback to the next model is a new job with its own key, which is what you want.
Which failures fall through
Not every error deserves another model. A 400 means this model rejected the request shape, such as a duration or aspect ratio outside its window, so the next model might accept it. A 429 or 502 is capacity, so moving on is reasonable. A 401 or 402 is about the account, and trying three more models only burns time.
| Response | Meaning | Chain action |
|---|---|---|
| 202 | Job accepted | Poll it |
| 400 unsupported_capability | Model cannot do this request | Next model |
| 429 rate_limited | Too many requests | Next model, or wait |
| 502 | Upstream problem | Next model |
| 401 / 402 | Bad key or low balance | Stop the chain |
| poll status failed | Render failed | Next model |
The code
Save it as chain.py, export SUME_API_KEY, and run it. All three models in the chain accept 5 seconds, 720p and 16:9, so one body works for every attempt. Check that before you add a model: minimax-h3, for example, has 480p and 768p rather than 720p, so it would need its own body.
import json, os, time, urllib.request, urllib.error
BASE = "https://api.sume.com/v1/videos"
CHAIN = ["gemini-omni-flash-1.1", "wan-3.0", "seedance-2"]
def call(method, url, body=None, key=None):
req = urllib.request.Request(url, method=method, data=body and json.dumps(body).encode())
req.add_header("Authorization", "Bearer " + os.environ["SUME_API_KEY"])
req.add_header("Content-Type", "application/json")
if key: req.add_header("Idempotency-Key", key)
try:
with urllib.request.urlopen(req, timeout=60) as r: return r.status, json.load(r)
except urllib.error.HTTPError as e: return e.code, json.load(e)
def render(prompt, row_id):
for model in CHAIN:
body = {"model": model, "prompt": prompt, "duration": 5,
"resolution": "720p", "aspect_ratio": "16:9"}
code, job = call("POST", BASE, body, f"{row_id}-{model}")
if code in (401, 402): raise SystemExit(job["error"])
if code != 202:
print(model, code, job["error"]["code"]); continue
while job.get("status") in ("pending", "in_progress"):
time.sleep(10)
code, job = call("GET", job["polling_url"])
if job.get("status") == "completed": return model, job["unsigned_urls"][0]
print(model, job.get("status"), job.get("error"))
raise RuntimeError("every model in the chain failed")
print(render("Slow dolly-in on a ceramic mug, steam rising", "row-42"))
Keep the chain short and honest
Fallbacks change the look of your output. A clip from Omni and a clip from Wan are different renders from the same prompt, and the audio differs too. Log the model that served each row, as the function returns it, so reviewers know which clips came from which family.
Order the chain by what you want, not by price. Put the model you tuned the prompt on first. The next two posts below show how to rank by cost when cost is the goal, and how to check a body against the catalog before you send it.
- Return the model name with the file, never just the URL.
- Cap the whole render with a wall-clock deadline, because three sequential polls can take a long time.
- Do not retry the same model after a failed status in this loop; the failed job already carries an error string worth reading.
Where to extend it
Add a callback_url on each submit and let the webhook trigger the next step if you do not want a worker sleeping. Add a per-model body builder if your chain includes a model with different resolutions or reference rules. Both are small changes to the same loop.
Sources
Related posts
More in Developers
- 'Video Router generate requires a catalog model id': the fix
The model must be an id from the catalog or an Auto alias. Provider names, vendor slugs and marketing names fail. List valid ids from /v1/videos/models first.
- wait_timeout_seconds 30 is not a 30-second video
The 30 in wait_timeout_seconds is how long your HTTP request may block, not how long a clip may run. A 30-second video job still needs a poll or webhook.
- waitForJob timeout in the Sume TypeScript SDK: keep the job id
waitForJob waits 20 minutes by default and throws SumeJobTimeoutError without cancelling the render. Catch it, store jobId, and resume later. Sume SDK 0.2.0.
- Wan 3.0 for 30 seconds in Node: submit, poll, save the MP4
A runnable Node 18+ script for Sume's /v1/videos: submit wan-3.0 at 30 seconds, poll with a deadline, stream the MP4 to disk. Price: $1.88 at 480p.
Written by Sume