Claude Message Batches 100,000 requests vs a Sume bulk run of 100

A Claude Message Batch holds 100,000 requests or 256 MB. A Sume bulk run queues 1 to 100 Format runs with a concurrency window of 1 to 16. Plan accordingly.

5 min readSume
All posts

A Claude Message Batch can hold up to 100,000 requests or 256 MB, whichever you hit first. A Sume bulk run holds 1 to 100 items, each a full Format run, with a concurrency of 1 to 16 that sets how many stay in flight. They are not the same kind of unit: Claude's requests are text completions, and Sume's items are video, image or avatar runs that take minutes each.

Claude's figures come from Claude's batch processing page. Sume's come from Bulk runs.

What are the limits on each side?

The Claude page says most batches finish within 1 hour, but results are available when all requests have finished or after 24 hours, whichever comes first, and a batch expires if processing is not done in 24 hours. Results stay downloadable for 29 days after creation.

Batch limits compared (read 2026-10-02)
ItemClaude Message BatchesSume bulk run
Max items100,000 requests or 256 MB100 items
ParallelismSet by the service and your rate limitsconcurrency 1 to 16, you choose
WindowExpires after 24 hoursNo queue expiry stated in the docs; each run has its own limits
Queue webhookNot covered hereNone; per-item communication.webhook_url only
List or cancel the queueCancel a batchNo list-queues or cancel-queue endpoint; cancel a child run
Per item unitOne Messages requestOne Format run: its own sandbox and agent turn

How do I move a 100,000-row job to Sume?

You cannot send it as one queue, because items is capped at 100 and a longer array is 400 invalid_request. Chunk the job into queues of up to 100, mint a fresh Idempotency-Key per chunk, and run them in sequence or in parallel within your plan's limits. Bulk runs also fix the number of in-flight children at concurrency, so extra chunks will not speed up work the workspace's generation concurrency already caps.

Ask first whether a Format run is the right unit. If the job is 100,000 text classifications, a Format is the wrong tool; Claude's batch API is built for that. A Format makes sense when each row needs generated media or a typed receipt.

What does a Sume chunk look like?

One chunk is one POST with an envelope of concurrency and items. The queue returns 202 and the first concurrency items are already running.

import os, uuid, requests

rows = [f"clip {i}" for i in range(1, 251)]
base = "https://api.sume.com/v1/formats/chase/product-promo/bulk-runs"
hdr = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
       "Content-Type": "application/json"}
queues = []
for i in range(0, len(rows), 100):
    body = {"concurrency": 8,
            "items": [{"instruction": r} for r in rows[i:i + 100]]}
    r = requests.post(base, json=body, timeout=30,
                      headers={**hdr, "Idempotency-Key": str(uuid.uuid4())})
    r.raise_for_status()
    queues.append(r.json()["data"]["id"])
print(queues)

What should I check when the queue finishes?

A queue is completed when every item is terminal, not when every item succeeded. Read counts.failed and counts.canceled, then read the reason on the child run at GET /v1/format-runs/{run_id}, because the queue item carries only a generic format_run_failed code. A failed item that never started a run has run_id: null and the create error on it.

The docs also say what Sume does not offer: no queue-level webhook, no list endpoint and no cancel endpoint for queues. If you need those, build them around the queue ids you stored.

Related reading: OpenAI's 50,000-request Batch API against a bulk run.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume