Claude Message Batches queue limits by tier vs Sume bulk runs
Claude's batch queue holds 200,000 requests on Start and up to 500,000 on Scale. A Sume bulk run takes 1 to 100 items, 16 at once. Different jobs.

A Claude batch queue holds 200,000 requests on the Start tier, 300,000 on Build and 500,000 on Scale, while one Sume bulk run takes 1 to 100 items with 1 to 16 running at once. They solve different jobs. Claude's batches are a discounted queue for text requests. A Sume bulk run is a window of paid media renders you watch.
What are the Claude numbers?
The batch processing page and the rate limits page give the figures.
| Topic | Start | Build | Scale |
|---|---|---|---|
| Batch processing queue | 200,000 | 300,000 | 500,000 |
| Batch requests per minute | 1,000 | 2,000 | 4,000 |
| Requests per batch | 100,000 | 100,000 | 100,000 |
What else does the batch page promise?
The page lists a 50% cost discount, a 256 MB size cap, expiry after 24 hours, and a note that most batches finish within an hour. Results come back per request as succeeded, errored, canceled or expired, and only succeeded requests are billed.
How does a Sume bulk run compare?
A bulk run creates up to 100 Format runs from one POST. Each item has the same body as a single run, and concurrency from 1 to 16 sets how many run at once. A malformed item returns 400 invalid_request with details.index before any queue exists, so you fix it before anything is billed.
Sume does not offer a batch discount that I could find in the docs, and the queue has no webhook, no list endpoint and no cancel-queue endpoint. You cancel a child run with POST /v1/format-runs/{run_id}/cancel.
- Creating a queue spends the write budget: Free 120, Pro 300, Startup 600, Scale 1200 per minute.
- Queue status is queued, running or completed. Completed means every item is terminal, not that every item succeeded.
- Branch on counts.failed and counts.canceled.
Which should I use?
For tens of thousands of text requests where a day's wait is fine, the Claude queue is built for it. For a hundred renders where each one has a spend cap and a receipt, use a bulk run, and split larger jobs into several queues with distinct idempotency keys. See the bulk runs reference.
Sources
Related posts
More in Comparisons
- Cloudflare Stream clip API vs Sume video trim: start and end
Cloudflare Stream clips by posting a source UID with startTimeSeconds and endTimeSeconds. Sume video-trim takes start plus end or duration, returning a new MP4.
- Cloudinary video 423 error on long videos vs Sume's 1800 s cap
Cloudinary returns 423 while a video over 30 or 60 minutes transforms asynchronously. Sume admits renders up to 1800 seconds and never returns 423.
- Cloudinary video trimming: percent or seconds vs Sume video trim
Cloudinary trims by start, end or duration, in seconds or a percentage of the asset. Sume video-trim takes seconds only: start plus one of end or duration.
- Creatify Boreal model_version on AI Avatar vs Sume quality tiers
Creatify added a boreal model_version to AI Avatar v1 and v2 in September 2026. In Sume you pick plus, standard or max quality. Compare the controls and limits.
Written by Sume