A fresh UUID per retry is not an idempotency key (Node, Sume)

Create the Idempotency-Key once outside the retry loop, retry only 429, 502, 503 and 504, and stop on 4xx. A Node submit helper for Sume /v1/videos.

5 min readSume
All posts

If your retry loop generates a new UUID on every attempt, it has no idempotency at all: each attempt is a new purchase. The key must be created once for the intent and reused by every retry. The Node helper below creates randomUUID() before the loop, retries only 429, 502, 503 and 504 plus network errors and timeouts, and throws on any other status. Sume's docs say a retry with the same key returns the original job instead of billing a second one.

Which statuses are worth retrying

Sume's error table separates problems you can fix from problems that go away. A 400 or 402 will fail again until you change the request or add funds. A 429 can be rate_limited, where you back off and use retry-after when it is present, or queue_full, where the workspace has no capacity left for another accepted job until one finishes. A 503 can be provider_capacity_exceeded, which the docs tell you to retry later with the same key. The 502 and 504 in the sample set are the helper's own choice for gateway errors and are not in Sume's table.

Retry decision for a submit (Sume docs, read 2026-10-09)
Status and codeRetry with the same key?Why
400 invalid_requestNoFix the body
402 insufficient_creditsNoAdd funds first
409 idempotency_conflictNoThe key was used with another payload
429 rate_limitedYes, after retry-afterRequest volume
429 queue_fullYes, once a job finishes or is canceledNo accepted-job capacity
503 provider_capacity_exceededYes, laterDispatch queue is full
Timeout or network errorYesYou do not know whether Sume received it

The helper

The helper sleeps 1, 2 and 4 seconds between three tries (1000 * 2 ** n milliseconds). Each fetch has a 20-second timeout through AbortSignal.timeout, and a TypeError or TimeoutError counts as retryable. Any other thrown error, including the one raised for a non-retryable status, leaves the loop at once.

When the loop gives up it says so in the message: a later retry with the same key is still safe. That is the point of keeping the key. Save it next to the request in your queue and the next worker run can pick it up.

import { randomUUID } from "node:crypto";

const RETRYABLE = new Set([429, 502, 503, 504]);

export async function submitWithRetry(body, tries = 3) {
  const key = randomUUID(); // once per intent, outside the loop
  for (let n = 0; n < tries; n++) {
    try {
      const res = await fetch("https://api.sume.com/v1/videos", {
        method: "POST",
        headers: {
          Authorization: `Bearer ${process.env.SUME_API_KEY}`,
          "Content-Type": "application/json",
          "Idempotency-Key": key,
        },
        body: JSON.stringify(body),
        signal: AbortSignal.timeout(20_000),
      });
      if (res.status === 202) return { key, job: await res.json() };
      if (!RETRYABLE.has(res.status)) throw new Error(`submit ${res.status}: ${await res.text()}`);
    } catch (err) {
      if (!(err instanceof TypeError || err.name === "TimeoutError")) throw err;
    }
    await new Promise((r) => setTimeout(r, 1000 * 2 ** n));
  }
  throw new Error("gave up; a later retry with the same key is still safe");
}

console.log(await submitWithRetry({ model: "minimax-h3", prompt: "Rain on a window", duration: 5 }));

Honor retry-after

The sample uses a fixed doubling schedule to stay short. For production, read the retry-after header on a 429 and wait at least that long instead of your own delay. The docs also tell you not to retry unsafe submit requests without an Idempotency-Key, which is why the key is created before the first attempt and not on the first failure.

For a worked cost: a 4-second seedance-2.5 clip at 480p is 4 x 0.268677 = $1.074708. A loop with a new key on each of three attempts that all reached Sume could create three jobs, about $3.22, while one key creates one job and the same $1.07. The code is the same length either way; the difference is one line above the loop.

  • Store the key with the job request, not in memory.
  • Cap total retry time, not only the count.
  • Alert on repeated queue_full; it means your plan's accepted-job capacity is too small for your batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume