Replicate API rate limits: 600 creates a minute, then 429
Replicate's API allows 600 prediction creates and 3,000 other requests per minute. Low credit and no card tighten it; over the limit you get a 429.

Replicate's API rate limits are 600 requests per minute for creating predictions and 3,000 requests per minute for every other endpoint. Short bursts above those numbers are allowed before throttling starts, limits get stricter as your credit runs low, and a request over the limit gets HTTP 429.
These facts come from Replicate's own Rate limits page, read on 2026-09-29. Replicate can change them; check the page before you size a job runner.
What are Replicate's API rate limits?
Replicate publishes two default limits and two conditions that lower them:
| Case | Limit |
|---|---|
| Creating predictions | 600 requests per minute |
| All other endpoints | 3,000 requests per minute |
| Short bursts | Allowed above the default limits before throttling |
| Credit running out | Stronger rate limits (no number published) |
| Granted credit, no payment method on file | 1 request per second, at most 6 requests per minute |
| Higher limits | Contact Replicate |
What happens when I hit the Replicate rate limit?
Replicate answers with status 429 and a body like { "detail" : "Request was throttled. Your rate limit resets in ~30s." }. The page doesn't document rate-limit headers, so read the detail text or back off on your own schedule.
In your client:
- Treat 429 as retryable: wait, then send the same request again, with growing delays and some jitter.
- Pace creates below 600 a minute with a client-side limiter rather than retrying into the wall. See client-side rate limiting in Python.
- Count status polls too: they are not prediction creates, so they fall under the 3,000-a-minute figure for other endpoints.
- Tell a 429 apart from a 5xx outage before retrying: see 429 vs 503.
Why am I throttled below 600 a minute?
Two account states lower the limit. First, as you approach running out of credit, Replicate applies stronger rate limits, to stop accidental overspending and to give you time to top up before you are shut off. Replicate's own advice is to set up credit auto-reload and keep your balance above $20. Second, if you've been granted credit and have no payment method on file, you're limited to 1 request per second with a maximum of 6 requests per minute. If a batch that normally runs fine starts getting 429s, check your balance first. For what Replicate charges, see Is the Replicate API free?.
How do I get higher Replicate API limits?
The page says to contact Replicate if you want higher limits. It doesn't publish a self-serve tier ladder, so there is no number to plan around above the defaults until Replicate agrees one with you.
Is a request limit the same as how many jobs run at once?
No. A rate limit caps how many HTTP requests you send in a window; a concurrency limit caps how many generations run at the same time. Replicate's rate-limits page covers only the first.
APIs split these differently, so check both numbers when you compare hosts. Sume, for example, gives each API key separate read and write request budgets and, separately, accepts jobs over its concurrency limit as `queued`; Sume API errors and rate limits has the numbers.
Sources
Related posts
More in Developers
- Runway API rate limit: usage tiers, concurrency and 429s
Runway's API has no requests-per-minute limit. Usage tiers cap concurrency per model, generations per 24 hours and monthly spend.
- C# speech to text: transcribe audio files with HttpClient
Speech to text in C#: POST the audio file's URL with HttpClient, poll the job, then read the transcript and word timestamps from the JSON result.
- Java subtitle generator API: burn captions with HttpClient
Generate subtitles from Java with the JDK HttpClient: POST the video URL to a captions API, poll the job, then read the captioned video_url.
- Talking avatar in JS: make one from Node, play it in React
A talking avatar in JavaScript: create it and send it a script from Node with the Sume SDK, wait for the job, then play the returned MP4 in React.
Written by Sume