Claude Message Batches queue limits by tier vs Sume bulk runs

Claude's batch queue holds 200,000 requests on Start and up to 500,000 on Scale. A Sume bulk run takes 1 to 100 items, 16 at once. Different jobs.

5 min readSume
All posts

A Claude batch queue holds 200,000 requests on the Start tier, 300,000 on Build and 500,000 on Scale, while one Sume bulk run takes 1 to 100 items with 1 to 16 running at once. They solve different jobs. Claude's batches are a discounted queue for text requests. A Sume bulk run is a window of paid media renders you watch.

What are the Claude numbers?

The batch processing page and the rate limits page give the figures.

Claude Message Batches limits (read 2026-10-02)
TopicStartBuildScale
Batch processing queue200,000300,000500,000
Batch requests per minute1,0002,0004,000
Requests per batch100,000100,000100,000

What else does the batch page promise?

The page lists a 50% cost discount, a 256 MB size cap, expiry after 24 hours, and a note that most batches finish within an hour. Results come back per request as succeeded, errored, canceled or expired, and only succeeded requests are billed.

How does a Sume bulk run compare?

A bulk run creates up to 100 Format runs from one POST. Each item has the same body as a single run, and concurrency from 1 to 16 sets how many run at once. A malformed item returns 400 invalid_request with details.index before any queue exists, so you fix it before anything is billed.

Sume does not offer a batch discount that I could find in the docs, and the queue has no webhook, no list endpoint and no cancel-queue endpoint. You cancel a child run with POST /v1/format-runs/{run_id}/cancel.

  • Creating a queue spends the write budget: Free 120, Pro 300, Startup 600, Scale 1200 per minute.
  • Queue status is queued, running or completed. Completed means every item is terminal, not that every item succeeded.
  • Branch on counts.failed and counts.canceled.

Which should I use?

For tens of thousands of text requests where a day's wait is fine, the Claude queue is built for it. For a hundred renders where each one has a spend cap and a receipt, use a bulk run, and split larger jobs into several queues with distinct idempotency keys. See the bulk runs reference.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume