429 on a chain step: retry with the same Idempotency-Key
A 429 on a trim or captions submit is rate_limited or queue_full. Wait for retry-after and resend the same body with the same Idempotency-Key.

When a trim or captions submit returns 429, wait for retry-after if it is present and resend the same body with the same Idempotency-Key. The two causes have different fixes, so read error.code first: rate_limited means your request volume went over the window, and queue_full means the workspace has no accepted generation capacity left. Neither one means your earlier steps were lost.
Which 429 is it
The docs list both under 429. The generation admission page says to retry queue_full with the same idempotency key after jobs finish or queued jobs are canceled, and to use retry-after for rate_limited.
| Code | Meaning | Do this |
|---|---|---|
rate_limited | Too many requests in the window | Wait retry-after, then retry with backoff and the same key |
queue_full | Concurrency and queue capacity are both full | Wait for jobs to finish or cancel queued ones, then retry with the same key |
402 insufficient_credits | Not a 429; credit admission failure | Top up; do not loop |
A retry loop for one step
Keep the key outside the loop. A new key per attempt is a new operation, and it could create a job if the first call did succeed but its response was lost. On 429, read retry-after for the delay. When it is absent, use your own growing delay with a ceiling. After a few attempts, store the step as pending with its key and return, so a restart can pick it up.
- Read
error.codebefore the delay: the response, not the status, tells you which limit you hit. - Resend the identical body with the identical key.
- Cap the attempts; a persistent
queue_fullis a signal to cancel work, not to wait forever. - Do not resubmit earlier steps. The render and trim jobs you already created are not affected.
Reads and writes are different buckets
A 429 names its budget in error.details.scope, which is read or write. Submitting a step is a write. Polling status_url is a read, and reads have a bucket 40 times larger on the shipped default, so a submit throttle does not stop your polling, and heavy polling does not use up the write budget. Check ratelimit-remaining instead of counting requests yourself.
Request rate is also not the same as generation capacity. Raising how fast you send does not raise the concurrency limit for your plan.
Backoff that does not hide the cause
Log the full error object, including error.code and, for a rate limit, error.details.scope. A write scope means that submits are the problem, so slow your submitter. A read scope means that polling is the problem, so lengthen the intervals you use and follow next_poll_after_seconds.
Do not retry in a tight loop. Honor retry-after, add jitter if many workers share the key, and keep a cap on attempts. For queue_full, think about the work in flight: if you do not need some queued jobs, cancel them before they start. That frees accepted capacity for the steps that matter.
The plan's concurrency is a separate limit from the request rate, so a faster sender does not get more renders running at once.
- Log
error.codeanddetails.scope. - Honor
retry-after. - Cancel unneeded queued jobs on
queue_full.
What not to change on a retry
Keep the body byte-for-byte the same, including optional fields. A changed field is a changed operation, and with the same key it returns 409 idempotency_conflict. Keep the key the same, because a new key could create a second paid job if the first request was accepted before the throttle response reached you.
Keep the earlier steps out of the retry. Only the step that got the 429 is resent. The chain state you stored tells you where you are, so the retry is one call and not three.
Sources
Related posts
More in Developers
- Brand name mispronounced by the AI voice? Four fixes and the cost
When a TTS voice says your brand name wrong, change the voice, the language, the spelling or the dictionary. Each retake on Sume costs 1 to 7 cents.
- Ack a Sume webhook in 10 s, then copy the 30-second MP4 later
Sume gives each delivery 10 seconds. Verify, store the job id, answer 204, and let a worker download the MP4 instead of doing it inside the handler.
- Ad run returns 400 invalid_attachment: count your 30 files first
A Format run carries at most 30 files, split 30 images, 10 videos, 10 audio. A Python counter for attachments plus media URLs inside input, before you call.
- Agency Black Friday: a team-owned Format needs a workspace key
A personal API key cannot run a Format that a team workspace owns. Sume answers 403 workspace_key_required, so issue the key inside the team before launch.
Written by Sume