Python fallback chain for Sume video models: 404 and 503 only

After the Sora API shutdown, a model chain must not retry everything. This Python function moves on at 404 and 503 and stops on 429, 402, 400, 409.

5 min readSume
All posts

A video model fallback chain should move to the next model on only two answers, 404 (the model id is not found) and 503 (provider capacity), and stop on everything else. A 429 means wait, a 402 means add funds, a 400 means your request is wrong, and a 409 means your key was reused. A different model fixes none of those. The Python below implements that rule and gives each model its own idempotency key.

The reason to build this now is the OpenAI deprecations page. It lists the Videos API and six Sora 2 ids as removed on 2026-09-24 and names no replacement.

What the vendor page says

Because the vendor lists no replacement, the pipeline has to pick one. A chain with an ordered model list is a way to make that choice a config value instead of a code change.

OpenAI deprecations page, Sora 2 and Videos API (read 2026-10-08)
ItemValue
Announced2026-03-24
Shutdown date2026-09-24
Listed idssora-2, sora-2-pro, sora-2-2025-10-06, sora-2-2025-12-08, sora-2-pro-2025-10-06, plus the Videos API
Recommended replacementNone listed

The chain

The key for each model is a hash of the job key, the model, and the prompt, cut to 32 characters. The same model and prompt always give the same key, so a retry of the whole function after a crash replays the accepted job instead of billing a second one. A different model gets a different key, which avoids a 409 idempotency_conflict caused by a changed payload.

import hashlib, json, os, urllib.error, urllib.request

BASE = os.environ.get("SUME_BASE", "https://api.sume.com")

def post(model, prompt, key):
    req = urllib.request.Request(
        BASE + "/v1/videos", method="POST",
        data=json.dumps({"model": model, "prompt": prompt}).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json", "Idempotency-Key": key})
    try:
        with urllib.request.urlopen(req, timeout=35) as r:
            return r.status, json.load(r)
    except urllib.error.HTTPError as e:
        return e.code, json.load(e)

def submit_with_fallback(models, prompt, job_key):
    for model in models:
        key = hashlib.sha256(f"{job_key}|{model}|{prompt}".encode()).hexdigest()[:32]
        status, body = post(model, prompt, key)
        if status == 202:
            return model, body
        if status not in (404, 503):  # 429, 402, 400, 409: switching models does not help
            raise RuntimeError(f"{status} {body.get('error', {}).get('code')} on {model}")
        print(f"{model}: {status}, trying next")
    raise RuntimeError("every model in the chain refused")

print(submit_with_fallback(["old-model", "ok-model"], "A mug on a desk", "order-8823"))

What the stub run showed

On a local stub, old-model answered 404 and ok-model answered 202. The function printed old-model: 404, trying next and returned the second model with its job envelope. A 429 on the first model would have raised at once, which is the point.

A 503 may clear in a few seconds. If you prefer, wait once and retry the same model before you move on; the docs say to retry later with the same key for provider_capacity_exceeded.

Before you trust the chain

  • Check each model's duration, resolution, and ratio limits against the catalog. A fallback that does not accept the same clip will fail with a 400.
  • Log which model served each job. Output differs between models, and your team should know which one it got.
  • Test the chain with a stub, not with paid jobs.

Putting the chain in config

Keep the model list outside the code, for example as a comma-separated environment variable, and split it at start. The first entry is the model you want; the rest are the ones you accept. A change of vendor, or a model that leaves the catalog, then needs a config edit and a restart, and not a deploy.

Order matters for cost as well as quality. Sume bills the provider list price times 1.25 for each model, so a fallback can cost more or less than the first choice. Read the price of each entry before you rely on it, and set a spend limit on the batch.

sume/auto is another option: you send it as the model, and Sume selects the family. It does not tell you which family ran, so use a pinned chain when you need to know.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume