Cloudflare Queues limits vs Sume queue_full 429: who retries?
Cloudflare Queues allows 100 retries and 24h delaySeconds. Sume answers 429 queue_full or rate_limited. How to back off in a consumer without duplicate jobs.

When a Cloudflare Queues consumer calls Sume and gets 429, let the queue do the waiting: retry the message with a delaySeconds value instead of sleeping inside the consumer. Cloudflare allows delaySeconds up to 24 hours, 100 retries and a 15 minute consumer wall time. Sume returns 429 queue_full or 429 rate_limited, and 402 insufficient_credits, and these three need different handling.
Limits on both sides
Values for Cloudflare are from its Queues limits page. Sume values are from its generation admission docs.
| Item | Cloudflare Queues | Sume |
|---|---|---|
| Message size | 128 KB | Request body per endpoint |
| Retries | Up to 100 | Use Idempotency-Key |
| Batch size | Up to 100 | Bulk runs up to 100 items |
| Delay | delaySeconds up to 24 h | retry-after on rate_limited |
| Consumer wall time | 15 min | Sync wait at most 30 s |
| Throughput and concurrency | 5,000 msg/s per queue | Plan only: Free 1, Pro 4, Startup 8, Scale 20 |
Handle the three 429 and 402 cases differently
rate_limited carries a retry-after hint, so use it as the delaySeconds. queue_full means your account's queue capacity, max(3, 5 times concurrency), is occupied. Retrying immediately only adds noise, so use a longer delay and fewer attempts. insufficient_credits is not transient: do not retry, alert a person.
Sume's concurrency is set by plan, so a Cloudflare queue that fans out wider than your concurrency will mostly produce waits. Treat the wave_size_hint in the generation_limits snapshot as a hint, not a guarantee.
- 429
rate_limited:message.retry({delaySeconds})using retry-after. - 429
queue_full: retry with a larger delay and cap attempts. - 402
insufficient_credits: ack the message and page someone. - Success 202: ack, and store the job id.
Keep retries from duplicating paid jobs
A queue redelivers on any consumer failure, including a crash after Sume accepted the request. Use the message id or your business id as the Idempotency-Key. A repeat with the same key and body returns the original job; a different body returns 409 idempotency_conflict, which tells you a producer bug changed the payload.
When to use bulk runs instead
If the messages are many items of one Format, one POST /v1/formats/{handle}/{slug}/bulk-runs takes up to 100 items with a concurrency setting and returns 202 with a queue you read at GET /v1/format-run-queues/{id}. A replayed key returns the old queue. Webhooks fire per item, not per queue.
Sources
Related posts
More in Developers
- Cloudflare Stream 200 MB upload limit: when you must use tus
Cloudflare Stream accepts basic uploads up to 200 MB and requires tus above that. Read the size from a Sume probe and route the file before you upload.
- Cloudflare Workflows step.sleep and a Sume job: poll without steps
Cloudflare Workflows step.sleep does not count toward the step limit, so a Sume job poll loop can wait cheaply. Limits, a loop shape and the retention catch.
- Codex 0.160 queued messages resume after reconnect: Sume safety
Codex 0.160 resumes unsent queued messages after a reconnect without duplicate sends. That covers messages, not tool effects, so keep Sume idempotency_key.
- Compare AI image models on the same prompt: a 20-line API script
MAI-Image-2.6 and Muse Image both claim No. 2 on Arena. Skip the leaderboard: run one prompt through several Sume image models and compare the URLs and cost.
Written by Sume