Check a Format output schema against Sume size limits first

Sume rejects schemas deeper than 10 levels, over 5000 properties or 120,000 characters. A local counter that catches max_depth and max_string_length early.

5 min readSume
All posts

Count your schema before you submit it. Sume's structured output docs set four size limits: nesting depth of 10 levels, 5,000 properties across the whole document, 1,000 values per enum, and 120,000 characters of total string length. Break one and the create fails with 400 output_schema_invalid and a violation named max_depth, max_properties, max_enum_values or max_string_length. A small local counter in your test suite finds the problem on a laptop instead of in a production request.

The limit most schemas hit is the last one

Depth and property counts are rarely a surprise. The string budget is. The docs define it as a document-wide sum over every property name, key and string value, so long description annotations on a large schema can exhaust it even when no single string is remarkable. Schemas generated from Pydantic models or from Zod with rich descriptions, plus enums of product categories, are the likeliest offenders.

The same page notes that $ref is limited to #/$defs/* and SumeMediaFile#, and that recursion through a named definition is fine. A self-referencing definition can still hit max_depth, since the depth limit applies to the document you send.

A counter you can run

The function below walks a schema and reports properties, characters, depth and the largest enum. It is an approximation of the server's rules, not a copy: it counts object and array nesting as levels and sums key and string lengths. Treat a result within 10 percent of a limit as a failure and trim. It needs only the Python standard library.

Add the real schema to a unit test that asserts each number stays under a comfortable margin, and your CI will catch the commit that pastes a 40,000-character taxonomy into a description.

import json

def measure(node, depth=1):
    props = chars = 0
    deepest = depth
    max_enum = 0
    if isinstance(node, dict):
        for k, v in node.items():
            chars += len(k)
            if k == "enum" and isinstance(v, list):
                max_enum = max(max_enum, len(v))
            if k == "properties" and isinstance(v, dict):
                props += len(v)
            p, c, d, e = measure(v, depth + (1 if k in ("properties", "items") else 0))
            props, chars, deepest, max_enum = props + p, chars + c, max(deepest, d), max(max_enum, e)
    elif isinstance(node, list):
        for v in node:
            p, c, d, e = measure(v, depth)
            props, chars, deepest, max_enum = props + p, chars + c, max(deepest, d), max(max_enum, e)
    elif isinstance(node, str):
        chars += len(node)
    return props, chars, deepest, max_enum

schema = {"type": "object", "additionalProperties": False, "required": ["title", "tags"],
          "properties": {"title": {"type": "string", "description": "Short title"},
                         "tags": {"type": "array", "items": {"type": "string", "enum": ["a", "b"]}}}}
props, chars, depth, enums = measure(schema)
print("properties", props, "of 5000; chars", chars, "of 120000; depth", depth, "of 10; max enum", enums, "of 1000")

What to trim first

Cut descriptions before structure. A field name and a type usually say enough, and instructions on how to fill the field belong in the Format or the run instruction, which is composed after the Format body and wins where they disagree. Move long enum lists into the instruction or into input. Flatten wrappers that only exist to group two fields.

If you need a very large output, ask for less per run. A bulk run of many small structured outputs is easier to validate and to retry than one schema that nears every limit.

read 2026-10-03
LimitValueViolation code
Nesting depth10 levelsmax_depth
Total properties5,000 across the documentmax_properties
Enum values1,000 per enummax_enum_values
Total string length120,000 charactersmax_string_length

Other rules to test alongside

Size is only one class of rejection. The root must be an object, every object needs additionalProperties: false, and every property must be listed in required, so optional fields become nullable type unions. oneOf, allOf, not and nullable: true are rejected; use anyOf and type unions. A single test that loads each Format's schema, runs the counter and checks those rules is cheap, and it fails before anything is reserved against your wallet.

Where limit errors surface

A schema that breaks a limit is rejected at create with 400 output_schema_invalid, so nothing runs and nothing is charged. The violation list names each { path, rule, message }, and the rule values match the names in the table. That makes the failure cheap but still annoying in production, which is why the local counter belongs in CI. If you generate schemas from code, run the counter on the generated output rather than on the source model, since defaults, descriptions and enum expansions only appear after generation.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume