Which Sume errors mean switch the video model, and which mean wait
Eleven documented Sume error codes and job categories sorted into three actions: fix the request, retry the same model later, or fall back to another model id.

Only a few Sume errors justify switching to a different video model. A 404 model_not_found means the id you sent is not in the catalog, so a fallback id is the right move. Most other errors, such as 429 queue_full, 503 provider_capacity_exceeded, or 402 insufficient_credits, say nothing about the model, and switching models would hide a problem that follows you to the next one.
The distinction matters when a vendor removes an API. OpenAI's deprecations page lists the Videos API and the sora-2 family with removal on 2026-09-24 and no replacement named (read 2026-10-08). A fallback chain built after that date should switch on the error that actually means the model is gone, not on every failure.
The decision table
Codes and meanings come from the errors page and the generation admission page. The third column is a recommendation for a fallback chain, not a Sume rule.
| Status and code | Documented meaning | Fallback chain action |
|---|---|---|
| 400 invalid_request | Body, model id shape, mode, or headers not valid | Fix the request. Do not switch. |
| 400 unsupported_parameter | A field the model does not accept, such as seed or size on /v1/videos | Drop the field. Switching hides the bug. |
| 401 unauthorized | Key missing or not valid | Stop. No model fixes this. |
| 402 insufficient_credits | Balance cannot cover the reservation | Stop, or try a cheaper request on the same model. |
| 404 model_not_found | The public model does not exist in this workspace | Switch to the next model id. |
| 409 idempotency_conflict | Same key, different payload | Use a new key for the new payload. |
| 429 rate_limited | Request volume over a protection limit | Wait for retry-after, same model, same key. |
| 429 queue_full | No accepted generation capacity left | Wait for jobs to finish. Same model, same key. |
| 503 provider_capacity_exceeded | Provider dispatch queue is full | Retry later with the same key. A switch is optional. |
| job error generation_rejected | Unsupported input | Correct the input, using the job events. |
| job error generation_unavailable | Generation is unavailable for now | Retry later. Then consider a switch. |
Why most errors should not trigger a switch
A queue-shaped error tells you about your workspace, not about the model. queue_full means the workspace has used all of its accepted generation capacity. Moving the same job to another model id would still land in the same queue. Sume's docs describe concurrency as a plan limit, with Free at 1 processing slot and 5 queue slots, which gives 6 accepted jobs.
A 402 is also workspace-wide. The reservation is the estimated cost of the request, so a cheaper model could pass where an expensive one failed. That is a billing decision, though, and your code should ask for it explicitly instead of making it silently.
A small fallback policy
Read retryable and next_action on a failed job before you resubmit. Those fields exist so that the client does not guess.
- Keep an ordered list of model ids in config, not in code.
- Switch only on
404 model_not_found, or on a failed job whose error category you have decided to treat as model-specific. - Send a new
Idempotency-Keywhen the payload changes, because a reused key with a different body returns409 idempotency_conflict. - Log the request id from the error envelope with every switch, so support can follow the path.
- Put
sume/autolast only if you accept that Sume does not disclose which family served the request.
Sources
Related posts
More in Developers
- How many Sume image rows list mask, background, quality, resolution
mask_url and background list on GPT Image 2.5 only, quality on three row types, resolution on a handful. A count of catalog rows by parameter, with a script.
- Which Sume MCP URL? mcp.sume.com/mcp, not the dev host or Studio
Add https://mcp.sume.com/mcp as the connector. The dev host is internal, and Studio Agent is a separate product, not the hosted MCP connector.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
Written by Sume