Does the Sume SDK retry POSTs? Only with an Idempotency-Key

createSumeClient retries 408, 429 and 5xx twice with backoff, but replays a POST only when it carries an Idempotency-Key. How to set it and tune maxRetries.

5 min readSume
All posts

Yes, but only when the request carries an Idempotency-Key. createSumeClient in @sume-com/sdk retries 408, 429, 5xx and transport failures, two times by default with exponential backoff and jitter, and it honours retry-after. A POST without an idempotency key is never replayed, because a replay would start and bill a second run.

This post is the exact rule, the two places a key comes from, and what to do when you submit a paid generation yourself instead of going through subscribeFormatRun. All of it is from the SDK run helpers page.

What does the client retry on its own?

The SDK documents two layers of retry, both on by default. The first is the client itself. The second is the wait loop in waitForRun, which tolerates consecutive failed status reads instead of throwing.

The table below is the shipped behaviour as the docs describe it for @sume-com/sdk@0.2.0.

SDK retry layers, read 2026-10-02
LayerDefaultOption
createSumeClient retries 408, 429, 5xx and transport failures2 retries, exponential backoff with jitter, honours retry-aftermaxRetries, timeout
waitForRun tolerates consecutive transient read failures6maxTransientFailures, onTransientError

Why is a POST retried only with an Idempotency-Key?

A retried GET is harmless. A retried POST /v1/video-1.0/generate after a dropped connection is not: the first request may have been accepted, and a second one would create and charge a second job. So the client replays a POST only if you gave it a key that lets the server recognise the repeat.

On a Format create the server side is documented in Calling a Format: the same key with the same body returns the original receipt with idempotency_hit: true and no second charge; the same key with a different body is 409 idempotency_conflict. A key is scoped to one Format and can be up to 255 characters.

The practical consequence is simple. Without a key, one network blip on a create is your problem, not the SDK's. With a key, the SDK can safely retry for you.

Where does the key come from?

For subscribeFormatRun the SDK generates a UUID unless you pass idempotencyKey yourself, or pass null to send none. A generated UUID protects against a transport retry inside one call, but not against your own process restarting and calling again, so derive the key from the thing being made, such as an order id plus a version you bump when you want a deliberate re-run.

For generation jobs there is no helper that adds one. You pass the header on the generated operation, as the docs show for generateVideoV1:

import { createSumeClient, generateVideoV1, waitForJob } from "@sume-com/sdk";

const client = createSumeClient({
  apiKey: process.env.SUME_API_KEY!,
  maxRetries: 3, // default is 2
});

const { data, error } = await generateVideoV1({
  client,
  headers: { "idempotency-key": "order-8823-hero-v1" },
  body: { prompt: "Slow push-in on a ceramic mug", mode: "async" },
});
if (error) throw new Error(JSON.stringify(error));

const job = await waitForJob(data!.data.request_id, { client });
console.log(job.status);

What still reaches your code as an error?

Generated operations do not throw on an API error. They resolve with { data, error, response }, so check error as above. The helpers subscribeFormatRun and waitForRun do throw, because a poll loop has nowhere to put a non-result. They throw SumeRunRequestError when the create is refused or a read fails and is not transient, and SumeRunTimeoutError when the timeout elapses.

Retries do not turn a client mistake into a success. A 4xx at create means nothing ran and nothing was charged, and the errors page notes that a failed create releases its key. Retrying a 403 insufficient_scope in a loop is called out there as the most common and most expensive mistake, so branch on code before you add retries of your own.

Read retryable and retry_after_seconds on the error envelope if you wrap the SDK in a queue worker. The SDK already waits for retry-after; do not stack a second backoff on top without a reason.

What does the SDK not do?

It does not retry a create that has no key, it does not cancel a run when a wait times out, and it does not give you an SSE stream. A SumeRunTimeoutError only means you stopped watching; the run keeps going and keeps spending, so store the run id and read it later with getFormatRunStatus, or call cancelFormatRun.

If you need a longer or shorter budget per request, set timeout on the client, and set maxRetries to 0 when you want to own retry policy entirely. Either way, keep the idempotency key stable across your own retries so the server collapses them.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume