API throttling vs rate limiting: what's the difference?
Rate limiting refuses requests over a quota with a 429; throttling slows or queues the excess instead. What each one means for your client.

Rate limiting caps how many requests a client may send in a time window and refuses the rest; the HTTP status for that refusal is 429 Too Many Requests, which may carry a Retry-After header. Throttling slows the excess down instead of refusing it: requests or jobs are delayed or queued and finish later. People often use the two words for the same thing, and many APIs do both, so read what the API does when you go over: an error you must retry, or a wait you must tolerate.
The 429 definition comes from RFC 6585. The worked example is Sume's API, which has both kinds of control, from Generation admission, Errors and rate limits and Authentication, all read on 2026-09-29.
What does a rate limit look like on the wire?
RFC 6585 defines 429 Too Many Requests as the user having sent too many requests in a given amount of time, and calls that "rate limiting". The response may include Retry-After to say how long to wait. The RFC leaves the counting to the server: per resource, per server, or across servers, keyed on credentials or a cookie. It also says servers are not required to send a 429: under attack, just dropping connections may be more appropriate.
So a rate limit is a refusal. Nothing happened, and the client decides when to try again.
What does throttling look like instead?
The request is accepted, but the work waits its turn. Sume's paid generation is an example. Concurrency is a dispatch limit, not a submit limit: when the workspace is at its concurrency limit, Sume still accepts valid jobs as queued while queue capacity remains, and workers move them to processing as slots free up. The client sees a slower job, not an error. The docs say not to treat queued as failure: store the job id and poll with backoff.
Throttling has a ceiling too. When the queue is also full, a new paid submit fails with 429 queue_full, a refusal that shares the status code with a rate limit but has a different cause.
How do rate limiting and throttling compare?
Side by side, the difference is what happens to the request over the limit. Sume's docs name four separate admission controls; only generation concurrency queues, while queue capacity, submit rate limits and balance refuse. Video job concurrency and queueing lists all four with per-plan numbers.
| Rate limiting | Throttling | |
|---|---|---|
| The request over the limit | Refused; nothing happens | Accepted, then delayed or queued |
| What the client sees | An error such as 429, maybe with Retry-After | A slower result, not an error |
| What the client does | Waits, then retries | Waits and polls for the result |
| Sume example | 429 rate_limited when a key's per-minute budget is spent | Paid jobs accepted as queued when generation concurrency is full |
| Its ceiling on Sume | Per-key budgets, reads and writes counted separately | Queue capacity; past it, 429 queue_full |
How should a client handle each one?
A rate limit protects the API from request volume; throttling protects the work behind it. Each answer needs its own response, below. 429 vs 503 covers the overload case, video job concurrency and queueing has Sume's per-plan numbers, and Python API rate limiting shows client-side pacing.
429 rate_limited: slow down. Sume gives each key a per-minute budget, with reads and writes counted separately; readratelimit-remainingand wait outretry-after. The429names the budget inerror.details.scope. Retry with backoff and an idempotency key. See rate limit headers.queued: nothing to fix. Poll the status with backoff until it is terminal, and don't resubmit because a local process timed out.429 queue_full: the docs say to wait for jobs to finish or cancel queued ones. Pacing requests doesn't clear it, and in current code a same-key retry replays the refusal, so resend with a new key once capacity is back.- Raising your request rate does not raise generation capacity; the plan's concurrency limit governs that separately.
Sources
Related posts
More in Developers
- API uptime SLA: what it promises, and Sume's terms
An API uptime SLA promises a monthly availability percentage and credits if missed. Sume's terms offer none; how to build around that.
- Automate video editing in Python with an editing API
Automate video editing in Python by calling an editing API with Requests: submit a caption, cut, or crop job, poll until it ends, chain the output.
- Bash for loop with curl: one API request per line
Loop over a file with while IFS= read -r, build each JSON body with jq --arg, send it with curl --fail-with-body, and pace it under the API's rate limit.
- Batch video processing by API: one edit, many videos
Batch video processing by API is a loop: one edit job per file, each with its own idempotency key, collected by webhook. How it works on Sume, and costs.
Written by Sume