Format run input 400: 64 top-level keys, 2 MiB, objects only

A Sume Format run rejects input that is not an object, has over 64 top-level keys, or exceeds 2 MiB. Group nested keys and check size before you send.

5 min readSume
All posts

The input on a Sume Format run must be a JSON object with at most 64 top-level keys and at most 2 MiB (2,097,152 UTF-8 bytes) on its compact serialization. An array, string or number is refused, and a violation is 400 invalid_request. Nested keys are not counted, so grouping related data under one key is free.

Those are the only structural checks. Sume publishes no field list for input: you choose the shape and the Format's recipe reads the keys it recognises. This post lists the checks, shows a pre-flight function, and separates them from the 4 MiB body cap and the media budget.

What does the API check on input?

The call page lists four checks and nothing else. Type, property count and size fail with a plain 400, and media references fail with 400 invalid_attachment because they share the run's attachment budget.

null and omission both mean no input, and an empty {} adds no file and no block at all, byte-identical to omitting the field. A body of {"input": {}} alone is 400 invalid_request, because a run must name at least one of instruction, input, previous_run_id or attachments.

Checks on input, read 2026-10-02
CheckRuleOn failure
TypeA JSON object; null or omitted means no input400
Property countAt most 64 top-level keys; nested keys not counted400
SizeAt most 2097152 UTF-8 bytes, compact serialization400
Media URLsHTTPS image, video or audio URLs at any depth share a budget of 30 files, 10 videos, 10 audio400 invalid_attachment

How do I stay under 64 keys?

Group. A flat object with one key per spreadsheet column hits 64 quickly on a wide sheet, while the same data under product, price, host and script is four keys. Only top-level keys count, so nesting is free, and the Format reads the nested keys it knows.

Grouping also reads better to the run. input is written whole to a file in the run's workspace and the agent is told it is caller data, not instructions, so a clear structure helps it find what it needs.

How is this different from the 4 MiB limit?

The 4 MiB cap applies to the whole request body and fails as 413 payload_too_large, as covered in the 413 post. The 2 MiB cap applies to input alone and fails as 400. Because input is capped lower, a body made mostly of input fails the 2 MiB check first.

Send media by URL rather than inline. Those URLs count toward the attachment budget, as explained in the media budget post.

A pre-flight check you can run

The function below applies the three structural checks before you call the API, and it measures the compact serialization the way the docs describe. Run it in your enqueue step, so a bad row fails in your own logs instead of costing a request. It does not count media URLs, which the API resolves itself.

import json

def check_input(value):
    if value is None:
        return "ok: no input"
    if not isinstance(value, dict):
        return "400: input must be a JSON object"
    if len(value) > 64:
        return f"400: {len(value)} top-level keys, max 64"
    size = len(json.dumps(value, separators=(",", ":"), ensure_ascii=False).encode())
    if size > 2097152:
        return f"400: {size} bytes, max 2097152"
    return f"ok: {len(value)} keys, {size} bytes"

print(check_input({"product": {"name": "Aurora"}, "price": {"list": "31,000"}}))
print(check_input([1, 2]))
print(check_input({f"k{i}": 1 for i in range(65)}))

Is input safe for untrusted text?

It is the right place for scraped copy and customer messages, rather than instruction, but the docs call it a trust boundary, not a sandbox. Runs are spend-capped, so a hostile payload's blast radius is bounded by the cap, but do not pass raw untrusted text on purpose. See the scraped copy post.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume