Idempotency key from the order id, not a fresh UUID per attempt
A random UUID generated inside the retry loop gives every attempt a new key and a new paid job. Derive the Sume Idempotency-Key from the order instead.

The Sume docs say to send Idempotency-Key on submit requests that you may retry after a client timeout or network failure, and to reuse a key only for the same operation and payload. A common bug is the opposite: the key is generated with crypto.randomUUID() inside the function that gets retried, so each attempt looks like a new operation and Sume creates and bills a new job.
The fix is to derive the key from something that identifies the business operation, not the attempt. An order id plus the asset name plus a version is enough, for example order-8823-hero-v1.
import { createSumeClient, generateVideoV1 } from "@sume-com/sdk";
const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! });
export async function submitHero(orderId: string, prompt: string) {
// Same order + same payload => same key => same job on every retry.
const key = `order-${orderId}-hero-v1`;
for (let attempt = 0; attempt < 3; attempt++) {
const { data, error } = await generateVideoV1({
client,
headers: { "idempotency-key": key },
body: { prompt, mode: "async" },
});
if (data) return data.data.request_id; // the job id
if (!isRetryable(error)) throw new Error(JSON.stringify(error));
await new Promise((r) => setTimeout(r, 2000 * 2 ** attempt));
}
throw new Error("submit failed after 3 attempts");
}
declare function isRetryable(error: unknown): boolean;
What the server does with the key
A retry with the same key and payload returns the original job. Reusing the key with a different operation or payload returns 409 idempotency_conflict; use a key again only for an exact retry. Agent Completions work the same way and mark a replay with idempotency_hit: true.
Failed creates matter too. On the Formats surface the docs say a failed create releases its key, so retrying after a 4xx that you fixed is a new attempt, not a replay. For queue_full and provider_capacity_exceeded, the docs say to retry later with the same key.
| Surface | Key | Replay signal |
|---|---|---|
| Generation jobs (REST) | Idempotency-Key header | Original job returned |
| Agent Completions | Idempotency-Key header | idempotency_hit: true |
| Format runs (SDK) | idempotencyKey option; auto UUID if omitted, null sends none | Finished run returns immediately |
| Hosted MCP writes and paid calls | idempotency_key field, required | Dedup, not approval |
Pitfalls
The SDK's subscribeFormatRun generates a UUID for you when you omit idempotencyKey. That is safe for a single call and unsafe for your own retry wrapper, because each outer retry gets a new default. Pass your own stable key if you wrap it.
- Change the version suffix when the prompt or options change, or you will get a
409. - Do not put secrets or personal data in the key; it appears in logs.
- A client timeout does not cancel the job. Keep the job id, poll it, and do not resubmit.
- One key per business operation, not one key per process.
A test you can run
Submit the same payload twice with the same key and compare the two job ids; they should be equal. Then submit it with a changed prompt and the same key and expect 409 idempotency_conflict. Finally, simulate a timeout by dropping the first response and retrying, and check that your usage ledger at GET /v1/usage shows one reservation, not two.
Keep keys short and readable. A key like order-8823-hero-v1 is easy to search in your logs and easy to bump to v2 when the creative changes.
Sources
Related posts
More in Developers
- Inspect transcript words without start or end: skip them in cuts
Inspect transcript words can omit start or end. Treat an untimed word as unknown, and never cut across a gap that contains one: 0 of its time is safe to remove.
- IVR phone menu prompts with TTS: 8 kHz mu-law WAV and cost per prompt
Make phone-menu prompts as 8 kHz mu-law or A-law WAV with Sume TTS. A 20-prompt menu bills 20 cents because each short job hits the 1-cent minimum.
- Sume job webhook retries: 370 s worst case, then poll or redeliver
Sume retries a job webhook 10 times, 30 s apart, with a 10 s timeout each: 270 s of gaps and up to 370 s in total. What to do when the window closes.
- jobs_wait any on 20 Wan 3.0 clips: the other 19 still bill $5.94
With wait_for any, jobs_wait returns when one of 20 Wan 3.0 480p 5-second clips ends, but the other 19 keep running and billing: $5.9375 of $6.25.
Written by Sume