OpenAI Batch 200 MB input file vs Sume's 4 MiB body: size your items

OpenAI's Batch API takes a 200 MB JSONL file. A Sume create body is capped at 4 MiB and input at 2 MiB, so media goes by URL and bulk items stay small.

5 min readSume
All posts

OpenAI's Batch API accepts an input file up to 200 MB and 50,000 requests per batch. Sume has no file upload for a bulk run: the whole create body is capped at 4 MiB (413 payload_too_large), input at 2 MiB and 64 top-level keys, and a queue holds at most 100 items. Send big media as URLs, and keep each item's JSON small.

OpenAI's limits are from its Batch API guide; Sume's are from Calling a Format and Bulk runs.

What are the size limits on each side?

OpenAI's guide also says only a 24h completion window is currently supported, output files are deleted 30 days after the batch completes, and each request needs a unique custom_id because results can arrive out of order.

Input size limits compared (read 2026-10-02)
LimitOpenAI Batch APISume Format run or bulk run
Whole payload200 MB input file4 MiB request body (413 payload_too_large)
Requests or items50,000 per batch1 to 100 per queue
Per-item dataPart of the request lineinput: 2 MiB and 64 top-level keys
How data arrivesUpload a JSONL fileOne JSON body per create
Item identityUnique custom_idPosition index, plus your own keys

What does Sume do with files and media?

Send media by URL. The docs say 413 payload_too_large fires over 4 MiB and the fix is to shrink input and send media by URL. Media URLs inside input share the run's attachment budget with attachments[], and a URL that cannot be fetched fails the create with 502 attachment_fetch_failed even though it is your input: its next_action is fix_input. Make the URL publicly reachable.

Because a bulk item is the same body as a single run, a 100-item queue could carry up to 100 inputs, but the 4 MiB cap applies to the whole create request. Do the arithmetic before you build a queue: 100 items at 2 MiB each would not fit in one body.

How do I check size before I submit?

Measure the compact serialization, which is how the docs count input.

import json

LIMIT_BODY = 4 * 1024 * 1024
LIMIT_INPUT = 2 * 1024 * 1024

def check(items: list[dict]) -> None:
    body = json.dumps({"concurrency": 4, "items": items},
                      separators=(",", ":")).encode()
    if len(body) > LIMIT_BODY:
        raise SystemExit(f"body {len(body)} bytes: split the queue")
    for i, it in enumerate(items):
        inp = json.dumps(it.get("input", {}), separators=(",", ":"))
        if len(inp.encode()) > LIMIT_INPUT or len(it.get("input", {})) > 64:
            raise SystemExit(f"item {i}: input too large or over 64 keys")

check([{"instruction": "clip 1", "input": {"url": "https://example.com/1.jpg"}}])
print("ok")

What if my job is bigger than 100 items?

Split it into several queues, each with its own Idempotency-Key, as in the bulk-run sizing post for OpenAI's 50,000-request limit. Sume does not offer a JSONL upload path in the docs, a queue-level webhook or a way to cancel a whole queue, so build those around stored queue ids.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume