A fresh UUID per retry is not an idempotency key (Node, Sume)
Create the Idempotency-Key once outside the retry loop, retry only 429, 502, 503 and 504, and stop on 4xx. A Node submit helper for Sume /v1/videos.

If your retry loop generates a new UUID on every attempt, it has no idempotency at all: each attempt is a new purchase. The key must be created once for the intent and reused by every retry. The Node helper below creates randomUUID() before the loop, retries only 429, 502, 503 and 504 plus network errors and timeouts, and throws on any other status. Sume's docs say a retry with the same key returns the original job instead of billing a second one.
Which statuses are worth retrying
Sume's error table separates problems you can fix from problems that go away. A 400 or 402 will fail again until you change the request or add funds. A 429 can be rate_limited, where you back off and use retry-after when it is present, or queue_full, where the workspace has no capacity left for another accepted job until one finishes. A 503 can be provider_capacity_exceeded, which the docs tell you to retry later with the same key. The 502 and 504 in the sample set are the helper's own choice for gateway errors and are not in Sume's table.
| Status and code | Retry with the same key? | Why |
|---|---|---|
| 400 invalid_request | No | Fix the body |
| 402 insufficient_credits | No | Add funds first |
| 409 idempotency_conflict | No | The key was used with another payload |
| 429 rate_limited | Yes, after retry-after | Request volume |
| 429 queue_full | Yes, once a job finishes or is canceled | No accepted-job capacity |
| 503 provider_capacity_exceeded | Yes, later | Dispatch queue is full |
| Timeout or network error | Yes | You do not know whether Sume received it |
The helper
The helper sleeps 1, 2 and 4 seconds between three tries (1000 * 2 ** n milliseconds). Each fetch has a 20-second timeout through AbortSignal.timeout, and a TypeError or TimeoutError counts as retryable. Any other thrown error, including the one raised for a non-retryable status, leaves the loop at once.
When the loop gives up it says so in the message: a later retry with the same key is still safe. That is the point of keeping the key. Save it next to the request in your queue and the next worker run can pick it up.
import { randomUUID } from "node:crypto";
const RETRYABLE = new Set([429, 502, 503, 504]);
export async function submitWithRetry(body, tries = 3) {
const key = randomUUID(); // once per intent, outside the loop
for (let n = 0; n < tries; n++) {
try {
const res = await fetch("https://api.sume.com/v1/videos", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify(body),
signal: AbortSignal.timeout(20_000),
});
if (res.status === 202) return { key, job: await res.json() };
if (!RETRYABLE.has(res.status)) throw new Error(`submit ${res.status}: ${await res.text()}`);
} catch (err) {
if (!(err instanceof TypeError || err.name === "TimeoutError")) throw err;
}
await new Promise((r) => setTimeout(r, 1000 * 2 ** n));
}
throw new Error("gave up; a later retry with the same key is still safe");
}
console.log(await submitWithRetry({ model: "minimax-h3", prompt: "Rain on a window", duration: 5 }));Honor retry-after
The sample uses a fixed doubling schedule to stay short. For production, read the retry-after header on a 429 and wait at least that long instead of your own delay. The docs also tell you not to retry unsafe submit requests without an Idempotency-Key, which is why the key is created before the first attempt and not on the first failure.
For a worked cost: a 4-second seedance-2.5 clip at 480p is 4 x 0.268677 = $1.074708. A loop with a new key on each of three attempts that all reached Sume could create three jobs, about $3.22, while one key creates one job and the same $1.07. The code is the same length either way; the difference is one line above the loop.
- Store the key with the job request, not in memory.
- Cap total retry time, not only the count.
- Alert on repeated queue_full; it means your plan's accepted-job capacity is too small for your batch.
Sources
Related posts
More in Developers
- Per-request timeout on Sume video polls: 15 s abort, then poll again
Give every poll GET its own 15 s AbortSignal.timeout so one hung read does not freeze a 30 s Wan or Seedance job loop. Reads per job per hour at 5, 10, 30 s.
- Agent Completion webhook retries: 10 attempts over about 3 hours
Sume tries an agent.run.terminal webhook up to 10 times. By the documented formula the nine waits add up to about 3 h 3 min, before jitter and Retry-After.
- AI video API billing units: per clip, per second, credits or tokens
Luma bills per generation, LTX per second, Runway and Vidu in credits, Google in tokens, MiniMax per second plus inputs, Sume in USD. Worked examples.
- Video fields Sume rejects: size, seed, provider.options, audio off
Sume's /v1/videos refuses size, seed and non-empty provider.options on every model, and Omni refuses generate_audio false. What to send in their place.
Written by Sume