Workers AI paid models 20/min, 50 with prepaid vs Sume plans
Cloudflare lifts paid Workers AI models from 20 to 50 requests per minute with prepaid credits. On Sume, top-ups do not raise concurrency; the plan does.

On Cloudflare, paying with prepaid credits raises the limit for paid Workers AI models from 20 to 50 requests per minute. On Sume, paying more does not raise throughput: generation concurrency is plan-only, and prepaid top-ups do not raise the processing concurrency limit.
Cloudflare's numbers are from its Workers AI limits page (last updated Sep 17, 2026) and Sume's from Generation admission, both read 2026-10-01.
What does Cloudflare's paid-model limit say?
The page says the limits apply per account, per model, for any model that requires the Workers Paid plan. Standard Workers AI billing gets 20 requests per minute; prepaid AI Gateway credits get 50. To receive the elevated limit you load prepaid AI Gateway credits and set the gateway's Workers AI billing setting accordingly.
What decides Sume throughput?
The plan. Processing concurrency is set per workspace by plan, with admin overrides possible, and queue capacity defaults to max(3, concurrency_limit x 5). The dashboard Concurrency tab is the source of truth; the table is a static guide.
| Plan | Processing | Queue (default) | Accepted |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
What happens when concurrency is full?
Nothing bad by itself: "Concurrency being full is not an error by itself." Valid jobs are accepted as queued while queue capacity remains, then move to processing. Only when the queue is also full does a submit fail with 429 queue_full. See concurrency full versus queue full.
So what does more money buy?
On Sume, a larger balance lets submits reserve their estimated cost; it does not widen the number of jobs processing at once. If you need more parallel work, check the plan row and the effective concurrency_limit reported in generation_limits, and queue the rest.
Sources
Related posts
More in Pricing
- Creatify Boreal price per second vs Sume's reserve ledger
Creatify prices Boreal at a cent per second of finished video. Sume shows each model's supported durations, reserves credits at admit and logs it in /v1/usage.
- ElevenLabs API pricing: USD, not credits, and Sume's wallet
ElevenLabs says API usage is billed in US dollars, not credits: TTS is $0.08 per 1,000 characters, $0.04 Flash/Turbo. Sume also bills a USD balance.
- ElevenLabs cost per extra minute: $0.36 to $0.17 by plan
ElevenLabs lists an extra minute at about $0.36 on Free, $0.20 Starter, $0.18 Creator and $0.17 Pro. Sume's TTS is billed per 1,000 characters, not minutes.
- ElevenLabs v4 3x credits on Creator+ until Oct 12 vs a USD wallet
ElevenLabs shows Eleven v4 with 3x credits on Creator+ until October 12. Sume has no credit multiplier: TTS is a USD price per transcript character.
Written by Sume