Which Sume API errors to retry and which to stop on: Node wrapper
A retry policy by error.code for Sume submits: retry rate_limited, queue_full, provider_capacity_exceeded; stop on 400, 401, 402, 409. Node wrapper with jitter.

Retry rate_limited, queue_full and provider_capacity_exceeded, always with the same Idempotency-Key. Stop and fix the request on 400, 401, 402 and 409. Treat provider_not_configured and other runtime errors as unavailable, and do not hammer them; the docs say the same.
Keying the decision on error.code rather than the status keeps the two 429s and the 503s apart. The table is from Errors and rate limits; the wrapper puts it into code.
Policy table
| Status | Code | Action |
|---|---|---|
| 400 | invalid_request | Fix the body, params or headers; no retry |
| 401 | unauthorized | Fix the key; no retry |
| 402 | insufficient_credits | Add funds or send a cheaper job |
| 409 | idempotency_conflict | Key reused for a different payload; mint a new key |
| 429 | rate_limited | Back off, honor retry-after |
| 429 | queue_full | Wait for capacity, same key |
| 503 | provider_capacity_exceeded | Retry later, same key |
| 503 | provider_not_configured | Do not retry aggressively; check the catalog |
The wrapper
Jitter spreads retries so a batch of clients does not return in lockstep.
const RETRY = new Set(["rate_limited", "queue_full", "provider_capacity_exceeded"]);
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
export async function submit(body, key, tries = 6) {
for (let n = 0; n < tries; n++) {
const res = await fetch("https://api.sume.com/v1/videos", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify(body),
});
if (res.ok) return res.json();
const err = (await res.json().catch(() => ({}))).error ?? {};
if (!RETRY.has(err.code)) throw new Error(err.code + " request_id=" + err.request_id);
const hinted = Number(res.headers.get("retry-after"));
const base = hinted > 0 ? hinted * 1000 : Math.min(60_000, 2 ** n * 1000);
await sleep(base + Math.random() * 500);
}
throw new Error("retries exhausted for key " + key);
}Job errors are different
A job that was accepted and then failed carries a public error with a category and a retryability hint. quota means add funds, generation_rejected means correct the input, and queue or generation_timeout means retry or poll later. Read those from the job, not from the submit response.
Logging that helps
Log the error code, the request id and the idempotency key on every failure, and nothing else that could be sensitive. Error bodies carry a request id that is safe to share with Sume support, and they are the fastest way to get a specific failure examined. Do not log API keys, signed URLs, raw media URLs or workspace identifiers.
- Alert on a rising share of
queue_full; it means you are submitting faster than the plan drains. - Alert on any
402, because it needs a human to add funds. - Treat repeated
provider_capacity_exceededas a signal to slow down, not to add retries.
Sources
Related posts
More in Developers
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
- Signed webhooks for Sume video runs: events, retries, verification
Sume sends one HMAC-SHA256 signed POST when a Format, Action, or Agent Completion run completes or fails. Verify the raw body and dedupe on request_id.
- Spend caps for unattended AI agents: how Sume bounds each run
An unattended agent has no one to approve spend, so Sume caps generation per run: required on Agent Completions, and up to $500 on Format runs.
Written by Sume