Does the Sume SDK retry POSTs? Only with an Idempotency-Key
createSumeClient retries 408, 429 and 5xx twice with backoff, but replays a POST only when it carries an Idempotency-Key. How to set it and tune maxRetries.

Yes, but only when the request carries an Idempotency-Key. createSumeClient in @sume-com/sdk retries 408, 429, 5xx and transport failures, two times by default with exponential backoff and jitter, and it honours retry-after. A POST without an idempotency key is never replayed, because a replay would start and bill a second run.
This post is the exact rule, the two places a key comes from, and what to do when you submit a paid generation yourself instead of going through subscribeFormatRun. All of it is from the SDK run helpers page.
What does the client retry on its own?
The SDK documents two layers of retry, both on by default. The first is the client itself. The second is the wait loop in waitForRun, which tolerates consecutive failed status reads instead of throwing.
The table below is the shipped behaviour as the docs describe it for @sume-com/sdk@0.2.0.
| Layer | Default | Option |
|---|---|---|
| createSumeClient retries 408, 429, 5xx and transport failures | 2 retries, exponential backoff with jitter, honours retry-after | maxRetries, timeout |
| waitForRun tolerates consecutive transient read failures | 6 | maxTransientFailures, onTransientError |
Why is a POST retried only with an Idempotency-Key?
A retried GET is harmless. A retried POST /v1/video-1.0/generate after a dropped connection is not: the first request may have been accepted, and a second one would create and charge a second job. So the client replays a POST only if you gave it a key that lets the server recognise the repeat.
On a Format create the server side is documented in Calling a Format: the same key with the same body returns the original receipt with idempotency_hit: true and no second charge; the same key with a different body is 409 idempotency_conflict. A key is scoped to one Format and can be up to 255 characters.
The practical consequence is simple. Without a key, one network blip on a create is your problem, not the SDK's. With a key, the SDK can safely retry for you.
Where does the key come from?
For subscribeFormatRun the SDK generates a UUID unless you pass idempotencyKey yourself, or pass null to send none. A generated UUID protects against a transport retry inside one call, but not against your own process restarting and calling again, so derive the key from the thing being made, such as an order id plus a version you bump when you want a deliberate re-run.
For generation jobs there is no helper that adds one. You pass the header on the generated operation, as the docs show for generateVideoV1:
import { createSumeClient, generateVideoV1, waitForJob } from "@sume-com/sdk";
const client = createSumeClient({
apiKey: process.env.SUME_API_KEY!,
maxRetries: 3, // default is 2
});
const { data, error } = await generateVideoV1({
client,
headers: { "idempotency-key": "order-8823-hero-v1" },
body: { prompt: "Slow push-in on a ceramic mug", mode: "async" },
});
if (error) throw new Error(JSON.stringify(error));
const job = await waitForJob(data!.data.request_id, { client });
console.log(job.status);What still reaches your code as an error?
Generated operations do not throw on an API error. They resolve with { data, error, response }, so check error as above. The helpers subscribeFormatRun and waitForRun do throw, because a poll loop has nowhere to put a non-result. They throw SumeRunRequestError when the create is refused or a read fails and is not transient, and SumeRunTimeoutError when the timeout elapses.
Retries do not turn a client mistake into a success. A 4xx at create means nothing ran and nothing was charged, and the errors page notes that a failed create releases its key. Retrying a 403 insufficient_scope in a loop is called out there as the most common and most expensive mistake, so branch on code before you add retries of your own.
Read retryable and retry_after_seconds on the error envelope if you wrap the SDK in a queue worker. The SDK already waits for retry-after; do not stack a second backoff on top without a reason.
What does the SDK not do?
It does not retry a create that has no key, it does not cancel a run when a wait times out, and it does not give you an SSE stream. A SumeRunTimeoutError only means you stopped watching; the run keeps going and keeps spending, so store the run id and read it later with getFormatRunStatus, or call cancelFormatRun.
If you need a longer or shorter budget per request, set timeout on the client, and set maxRetries to 0 when you want to own retry policy entirely. Either way, keep the idempotency key stable across your own retries so the server collapses them.
Sources
Related posts
More in Developers
- @sume-com/sdk waitForJob is not exported: a 26-line replacement
The docs show waitForJob, but the npm build of @sume-com/sdk 0.2.0 does not export it. Here is a fetch version with 429 tolerance and a deadline.
- Check an Avatar payload with sume tools schema before --confirm-paid
sume tools schema avatar-videos.create --json prints the exact fields before you pass --confirm-paid. A read-only pre-flight loop for a spending CLI command.
- Cost of one Sume agent thread or turn: /v1/usage thread_id and job_id
Pass thread_id to GET /v1/usage to sum one Studio Agent thread, or a turn's job id as job_id for that turn plus every job it commissioned. Fields and a request.
- Sume /v1/balance: next_expires_at and the expiring-soon fields
GET /v1/balance returns USD micros and cents, a funded or empty state, and an expiration block with the next expiry and the amount expiring soon. Field list.
Written by Sume