AI video API fallback: retry on another model when a job fails

Chain seedance-2.5, seedance-2 and kling-3 on Sume: poll status_url, read the job error category, and resubmit the brief to the next model.

5 min readSume
All posts

To fall back to another video model in the Sume API, submit the brief to your first choice with POST /v1/video-router/generate, poll the job's status_url until terminal is true, and if sume_status is failed, read the job's error and submit the same brief to the next model in a short chain, with a new Idempotency-Key for each model. Only chain models whose limits cover your request, because a model that cannot accept your duration, resolution or aspect ratio will fail at submit instead of rescuing the clip.

This post builds that chain for a 9:16, 720p, 8-second clip across seedance-2.5, seedance-2 and kling-3. The envelopes come from the Video Router docs and the Video generation docs; the failure vocabulary comes from Errors and rate limits, all read on 2026-10-03.

Which models can share one request?

A fallback is only useful if the second model accepts the first model's request unchanged. Send the shared subset and nothing model-specific. For a vertical 8-second clip that subset is 720p, 9:16 and 8 seconds, and every model in the table covers it.

Leave generate_audio out of the shared body. Its default is each model's own audio capability, so the fallback clip may come back with a different audio setup than the clip you lost. Treat that as part of the trade rather than a bug.

Limits of a three-model fallback chain, from Sume's Video Router and Video generation docs, read 2026-10-03
Model idDurationResolutionAspect ratiosRole in the chain
seedance-2.54-30 s480p, 720p, 1080p21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus autoFirst choice
seedance-24-15 s480p, 720p, 1080p21:9, 16:9, 4:3, 1:1, 3:4, 9:16Same family, shorter ceiling
kling-34-15 s720p, 1080p16:9, 9:16, 1:1Different model family as a last resort

Which failures deserve a fallback?

A failed job carries public error metadata: category, stage, retryable, retry_after_seconds, public_reason and next_action. Read it from GET /v1/jobs/{id}, where it sits at data.job.error. Use the category to decide, rather than treating every failure alike.

A validation failure means your input was wrong, so the next model will probably fail the same way. quota and auth mean the workspace, not the model, is the problem, so the chain should stop. generation_unavailable, generation_timeout, runtime_unavailable and worker_timeout are the cases where moving on is cheap and sensible. generation_rejected is the ambiguous one: Sume's docs say to inspect events and fix unsupported input, so read GET /v1/jobs/{id}/events before you assume another model will accept the same media.

  • Stop the chain on auth or quota.
  • Fix the request and do not fall back on validation.
  • Fall back, or retry later, on generation_unavailable, generation_timeout, runtime_unavailable and worker_timeout.
  • On generation_rejected, read the job events first.
  • Never resubmit just because your own process timed out; the job may still be running, so poll its status_url instead.

What does the chain look like in Python?

The script below submits to each model in turn, polls with the interval Sume suggests in next_poll_after_seconds, and stops at the first Sume-hosted artifact URL, or at an auth or quota failure. Each attempt gets its own idempotency key built from your run id and the model id, so a retry of the same attempt after a network error adopts the original job, while a different model is a genuinely new request.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
BRIEF = {"prompt": "A ceramic mug on a desk, steam rising", "resolution": "720p",
         "duration": 8, "aspect_ratio": "9:16", "mode": "async"}

def attempt(model, run_id):
    r = requests.post(f"{API}/v1/video-router/generate", timeout=60,
                      headers={**H, "Idempotency-Key": f"{run_id}-{model}"},
                      json={**BRIEF, "model": model})
    if not r.ok:
        return None, {"http": r.status_code}
    env = r.json()["data"]
    while True:
        s = requests.get(env["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(s.get("next_poll_after_seconds") or 5)
    if s["sume_status"] != "completed":
        j = requests.get(f"{API}/v1/jobs/{env['job']['id']}", headers=H, timeout=30)
        return None, j.json()["data"]["job"].get("error")
    res = requests.get(env["result_url"], headers=H, timeout=30).json()["data"]["result"]
    return res["artifacts"][0]["url"], None

for model in ["seedance-2.5", "seedance-2", "kling-3"]:
    url, err = attempt(model, "mug-clip-001")
    print(model, url or err)
    if url or (err or {}).get("category") in ("auth", "quota"):
        break

What changes when the fallback model answers?

The clip is a new take from a different model, so expect a different look. Sume's docs list no seed field on these models, so even the same model cannot be asked to repeat a take. Record job.model next to every delivered clip so your team knows which model made it, and keep the fallback chain short enough that a reviewer can tell the models apart.

Pinned models and sume/auto solve different problems. The chain above is for a pipeline that must name the model on every clip. If you would rather let Sume choose and do not need to know which model ran, pin a model or use sume/auto compares the two. Billing for a failed or canceled job is covered separately in Charged for a failed Seedance or Kling job?.

Where does the chain stop being safe?

Do not add models whose input shape differs from the brief. grok-imagine-video-1.5 needs an image_url first frame, and gemini-omni-flash-1.1 accepts only 3 to 10 seconds in 16:9 or 9:16 with audio always on, so each would need its own request body. If your brief has references, check which models take them before you chain: kling-3 rejects reference_*_urls entirely.

Poll from your own process and keep the wait in your client. Sume's sync mode only blocks up to 30 seconds, which is shorter than a typical video job, and a timeout there is not a job outcome.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume