429 after a traffic spike: ramp up generation jobs in waves
Anthropic warns a sharp usage increase can trigger 429 acceleration limits. For Sume jobs, size each wave from generation_limits and the in-flight budget.

If a 429 shows up right after a jump in traffic, ramp up in waves instead of firing everything at once. Anthropic's page says a sharp increase in usage can trigger acceleration limits and advises ramping gradually. For Sume generation jobs, size each wave from the generation_limits snapshot on submit responses and refresh it before the next wave.
Anthropic facts are from its Rate limits page (listed under Sources) and Sume facts from Generation admission, read 2026-09-30.
What does Anthropic say about sudden increases?
Its note says you might get 429 errors because of acceleration limits if your organization has a sharp increase in usage, and to avoid them you should ramp up traffic gradually and maintain consistent usage patterns. Sume's docs describe no equivalent acceleration limit; its 429s are rate_limited and queue_full.
How do I size a wave of Sume jobs?
The docs give a budget for new in-flight work: max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. wave_size_hint is only a submission hint and is never a concurrency limit. At zero headroom, wait and refresh before submitting more.
| Snapshot field | Value |
|---|---|
| concurrency_limit | 100 |
| queued_jobs_limit | 500 |
| queue_capacity_remaining (no jobs running) | 600 |
| wave_size_hint | 450 |
| New in-flight budget with 30 processing, 10 queued | 60 |
What if I get queue_full anyway?
queue_full means every accepted slot is used. Stop adding work, poll existing jobs until one is terminal, cancel queued jobs you no longer need, and retry with the same idempotency key once capacity opens. Concurrency being full alone is not an error: valid jobs are accepted as queued.
Does a slower ramp change cost?
No. Submit pacing decides when jobs start, not what they cost; the estimate is reserved at submit and captured on success. Do not resubmit a paid request because a local worker timed out; reuse the Idempotency-Key for retries of the same intent.
Sources
Related posts
More in Developers
- AI SDK isLoopFinished vs the 20-step cap for Sume video jobs
ToolLoopAgent stops after 20 steps by default. How many jobs_wait calls a Sume video render needs, and the per-call guards to keep if you lift the cap.
- AI SDK MCP client close() in onEnd: Sume jobs keep running
Closing the AI SDK MCP client in onEnd ends the tool connection, not the Sume jobs already submitted. Track job ids, then read or cancel them.
- Airtable script 50-fetch limit: send 100 records as one bulk run
An Airtable automation script gets 50 fetches and a 30-second timeout. One Sume bulk-run request carries up to 100 items, so one fetch beats a per-record loop.
- Airtable has no webhook signature check: verify Sume first
Airtable's When webhook received trigger cannot verify signatures and caps payloads at 100kb. Verify the Sume signature in a relay and forward a small object.
Written by Sume