Video model fallback chain in Python with one key per model

A Python chain that tries Gemini Omni, Wan 3.0 and Seedance 2 on Sume in order, stops on 401 or 402, and uses a separate idempotency key for each model.

5 min readSume
All posts

Write the fallback as a loop over catalog model ids, send each attempt with its own Idempotency-Key, and stop the whole chain on 401 or 402, because every later model would fail the same way. The Python below tries gemini-omni-flash-1.1, then wan-3.0, then seedance-2, and returns the first completed file.

This matters now because OpenAI's deprecations page lists the Sora video models and the Videos API as removed on 2026-09-24 with no replacement named. A pipeline that had one video vendor has one video vendor fewer, and the lesson is to never hard-wire a single model id again.

Why the key is per model

Sume's idempotency rule is that a replay with the same key and the same body returns the original job, while the same key with a different body returns 409 conflict. A different model is a different body. If you reuse one key across the chain, the second attempt collides with the first and you get a 409 instead of a render.

Building the key as the row id plus the model id fixes both sides. A crash and restart replays each attempt to its original job, so you never pay twice for the same row and model. A fallback to the next model is a new job with its own key, which is what you want.

Which failures fall through

Not every error deserves another model. A 400 means this model rejected the request shape, such as a duration or aspect ratio outside its window, so the next model might accept it. A 429 or 502 is capacity, so moving on is reasonable. A 401 or 402 is about the account, and trying three more models only burns time.

Fallback decision per response (Sume docs, read 2026-10-05)
ResponseMeaningChain action
202Job acceptedPoll it
400 unsupported_capabilityModel cannot do this requestNext model
429 rate_limitedToo many requestsNext model, or wait
502Upstream problemNext model
401 / 402Bad key or low balanceStop the chain
poll status failedRender failedNext model

The code

Save it as chain.py, export SUME_API_KEY, and run it. All three models in the chain accept 5 seconds, 720p and 16:9, so one body works for every attempt. Check that before you add a model: minimax-h3, for example, has 480p and 768p rather than 720p, so it would need its own body.

import json, os, time, urllib.request, urllib.error
BASE = "https://api.sume.com/v1/videos"
CHAIN = ["gemini-omni-flash-1.1", "wan-3.0", "seedance-2"]

def call(method, url, body=None, key=None):
    req = urllib.request.Request(url, method=method, data=body and json.dumps(body).encode())
    req.add_header("Authorization", "Bearer " + os.environ["SUME_API_KEY"])
    req.add_header("Content-Type", "application/json")
    if key: req.add_header("Idempotency-Key", key)
    try:
        with urllib.request.urlopen(req, timeout=60) as r: return r.status, json.load(r)
    except urllib.error.HTTPError as e: return e.code, json.load(e)

def render(prompt, row_id):
    for model in CHAIN:
        body = {"model": model, "prompt": prompt, "duration": 5,
                "resolution": "720p", "aspect_ratio": "16:9"}
        code, job = call("POST", BASE, body, f"{row_id}-{model}")
        if code in (401, 402): raise SystemExit(job["error"])
        if code != 202:
            print(model, code, job["error"]["code"]); continue
        while job.get("status") in ("pending", "in_progress"):
            time.sleep(10)
            code, job = call("GET", job["polling_url"])
        if job.get("status") == "completed": return model, job["unsigned_urls"][0]
        print(model, job.get("status"), job.get("error"))
    raise RuntimeError("every model in the chain failed")

print(render("Slow dolly-in on a ceramic mug, steam rising", "row-42"))

Keep the chain short and honest

Fallbacks change the look of your output. A clip from Omni and a clip from Wan are different renders from the same prompt, and the audio differs too. Log the model that served each row, as the function returns it, so reviewers know which clips came from which family.

Order the chain by what you want, not by price. Put the model you tuned the prompt on first. The next two posts below show how to rank by cost when cost is the goal, and how to check a body against the catalog before you send it.

  • Return the model name with the file, never just the URL.
  • Cap the whole render with a wall-clock deadline, because three sequential polls can take a long time.
  • Do not retry the same model after a failed status in this loop; the failed job already carries an error string worth reading.

Where to extend it

Add a callback_url on each submit and let the webhook trigger the next step if you do not want a worker sleeping. Add a per-model body builder if your chain includes a model with different resolutions or reference rules. Both are small changes to the same loop.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume