Which Sume errors mean switch the video model, and which mean wait

Eleven documented Sume error codes and job categories sorted into three actions: fix the request, retry the same model later, or fall back to another model id.

5 min readSume
All posts

Only a few Sume errors justify switching to a different video model. A 404 model_not_found means the id you sent is not in the catalog, so a fallback id is the right move. Most other errors, such as 429 queue_full, 503 provider_capacity_exceeded, or 402 insufficient_credits, say nothing about the model, and switching models would hide a problem that follows you to the next one.

The distinction matters when a vendor removes an API. OpenAI's deprecations page lists the Videos API and the sora-2 family with removal on 2026-09-24 and no replacement named (read 2026-10-08). A fallback chain built after that date should switch on the error that actually means the model is gone, not on every failure.

The decision table

Codes and meanings come from the errors page and the generation admission page. The third column is a recommendation for a fallback chain, not a Sume rule.

Documented Sume error codes and the suggested fallback action (docs read 2026-10-08)
Status and codeDocumented meaningFallback chain action
400 invalid_requestBody, model id shape, mode, or headers not validFix the request. Do not switch.
400 unsupported_parameterA field the model does not accept, such as seed or size on /v1/videosDrop the field. Switching hides the bug.
401 unauthorizedKey missing or not validStop. No model fixes this.
402 insufficient_creditsBalance cannot cover the reservationStop, or try a cheaper request on the same model.
404 model_not_foundThe public model does not exist in this workspaceSwitch to the next model id.
409 idempotency_conflictSame key, different payloadUse a new key for the new payload.
429 rate_limitedRequest volume over a protection limitWait for retry-after, same model, same key.
429 queue_fullNo accepted generation capacity leftWait for jobs to finish. Same model, same key.
503 provider_capacity_exceededProvider dispatch queue is fullRetry later with the same key. A switch is optional.
job error generation_rejectedUnsupported inputCorrect the input, using the job events.
job error generation_unavailableGeneration is unavailable for nowRetry later. Then consider a switch.

Why most errors should not trigger a switch

A queue-shaped error tells you about your workspace, not about the model. queue_full means the workspace has used all of its accepted generation capacity. Moving the same job to another model id would still land in the same queue. Sume's docs describe concurrency as a plan limit, with Free at 1 processing slot and 5 queue slots, which gives 6 accepted jobs.

A 402 is also workspace-wide. The reservation is the estimated cost of the request, so a cheaper model could pass where an expensive one failed. That is a billing decision, though, and your code should ask for it explicitly instead of making it silently.

A small fallback policy

Read retryable and next_action on a failed job before you resubmit. Those fields exist so that the client does not guess.

  • Keep an ordered list of model ids in config, not in code.
  • Switch only on 404 model_not_found, or on a failed job whose error category you have decided to treat as model-specific.
  • Send a new Idempotency-Key when the payload changes, because a reused key with a different body returns 409 idempotency_conflict.
  • Log the request id from the error envelope with every switch, so support can follow the path.
  • Put sume/auto last only if you accept that Sume does not disclose which family served the request.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume