AI video API fallback: retry on another model when a job fails
Chain seedance-2.5, seedance-2 and kling-3 on Sume: poll status_url, read the job error category, and resubmit the brief to the next model.

To fall back to another video model in the Sume API, submit the brief to your first choice with POST /v1/video-router/generate, poll the job's status_url until terminal is true, and if sume_status is failed, read the job's error and submit the same brief to the next model in a short chain, with a new Idempotency-Key for each model. Only chain models whose limits cover your request, because a model that cannot accept your duration, resolution or aspect ratio will fail at submit instead of rescuing the clip.
This post builds that chain for a 9:16, 720p, 8-second clip across seedance-2.5, seedance-2 and kling-3. The envelopes come from the Video Router docs and the Video generation docs; the failure vocabulary comes from Errors and rate limits, all read on 2026-10-03.
Which models can share one request?
A fallback is only useful if the second model accepts the first model's request unchanged. Send the shared subset and nothing model-specific. For a vertical 8-second clip that subset is 720p, 9:16 and 8 seconds, and every model in the table covers it.
Leave generate_audio out of the shared body. Its default is each model's own audio capability, so the fallback clip may come back with a different audio setup than the clip you lost. Treat that as part of the trade rather than a bug.
| Model id | Duration | Resolution | Aspect ratios | Role in the chain |
|---|---|---|---|---|
seedance-2.5 | 4-30 s | 480p, 720p, 1080p | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus auto | First choice |
seedance-2 | 4-15 s | 480p, 720p, 1080p | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | Same family, shorter ceiling |
kling-3 | 4-15 s | 720p, 1080p | 16:9, 9:16, 1:1 | Different model family as a last resort |
Which failures deserve a fallback?
A failed job carries public error metadata: category, stage, retryable, retry_after_seconds, public_reason and next_action. Read it from GET /v1/jobs/{id}, where it sits at data.job.error. Use the category to decide, rather than treating every failure alike.
A validation failure means your input was wrong, so the next model will probably fail the same way. quota and auth mean the workspace, not the model, is the problem, so the chain should stop. generation_unavailable, generation_timeout, runtime_unavailable and worker_timeout are the cases where moving on is cheap and sensible. generation_rejected is the ambiguous one: Sume's docs say to inspect events and fix unsupported input, so read GET /v1/jobs/{id}/events before you assume another model will accept the same media.
- Stop the chain on
authorquota. - Fix the request and do not fall back on
validation. - Fall back, or retry later, on
generation_unavailable,generation_timeout,runtime_unavailableandworker_timeout. - On
generation_rejected, read the job events first. - Never resubmit just because your own process timed out; the job may still be running, so poll its
status_urlinstead.
What does the chain look like in Python?
The script below submits to each model in turn, polls with the interval Sume suggests in next_poll_after_seconds, and stops at the first Sume-hosted artifact URL, or at an auth or quota failure. Each attempt gets its own idempotency key built from your run id and the model id, so a retry of the same attempt after a network error adopts the original job, while a different model is a genuinely new request.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
BRIEF = {"prompt": "A ceramic mug on a desk, steam rising", "resolution": "720p",
"duration": 8, "aspect_ratio": "9:16", "mode": "async"}
def attempt(model, run_id):
r = requests.post(f"{API}/v1/video-router/generate", timeout=60,
headers={**H, "Idempotency-Key": f"{run_id}-{model}"},
json={**BRIEF, "model": model})
if not r.ok:
return None, {"http": r.status_code}
env = r.json()["data"]
while True:
s = requests.get(env["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(s.get("next_poll_after_seconds") or 5)
if s["sume_status"] != "completed":
j = requests.get(f"{API}/v1/jobs/{env['job']['id']}", headers=H, timeout=30)
return None, j.json()["data"]["job"].get("error")
res = requests.get(env["result_url"], headers=H, timeout=30).json()["data"]["result"]
return res["artifacts"][0]["url"], None
for model in ["seedance-2.5", "seedance-2", "kling-3"]:
url, err = attempt(model, "mug-clip-001")
print(model, url or err)
if url or (err or {}).get("category") in ("auth", "quota"):
breakWhat changes when the fallback model answers?
The clip is a new take from a different model, so expect a different look. Sume's docs list no seed field on these models, so even the same model cannot be asked to repeat a take. Record job.model next to every delivered clip so your team knows which model made it, and keep the fallback chain short enough that a reviewer can tell the models apart.
Pinned models and sume/auto solve different problems. The chain above is for a pipeline that must name the model on every clip. If you would rather let Sume choose and do not need to know which model ran, pin a model or use sume/auto compares the two. Billing for a failed or canceled job is covered separately in Charged for a failed Seedance or Kling job?.
Where does the chain stop being safe?
Do not add models whose input shape differs from the brief. grok-imagine-video-1.5 needs an image_url first frame, and gemini-omni-flash-1.1 accepts only 3 to 10 seconds in 16:9 or 9:16 with audio always on, so each would need its own request body. If your brief has references, check which models take them before you chain: kling-3 rejects reference_*_urls entirely.
Poll from your own process and keep the wait in your client. Sume's sync mode only blocks up to 30 seconds, which is shorter than a typical video job, and a timeout there is not a job outcome.
Sources
Related posts
More in Developers
- Virtual try-on API: which Sume call returns an image, which a video
Need a try-on photo or a try-on clip? On Sume the two catalog try-on Formats return video; a still comes from the image API. The table, plus one call for each.
- Test a webhook endpoint before go-live: a Sume CI gate (Python)
Use POST /v1/webhooks/test-deliveries to fire a signed webhook.test at your deployed URL and fail the deploy unless it answers 2xx. Python script included.
- Zod 4 discriminated union for Sume job and run webhooks (TypeScript)
Parse Sume job.* and format.run.terminal webhooks with one Zod 4 discriminatedUnion: typed branches, degraded runs, oversized receipts. Tested with Zod 4.
- Zod 4 toJSONSchema to Sume output_schema: nullable, not optional
z.toJSONSchema works for a Sume Format output_schema if you use nullable instead of optional. A tested table of what passes and what the validator rejects.
Written by Sume