Sume output schema limits: depth 10, 5,000 properties, 120,000 chars

How big can a Sume output_schema be? Depth 10, 5,000 properties and 120,000 characters, and a refusal before any spend. What to flatten and where it fails.

4 min readSume
All posts

A Sume output schema may nest ten levels deep, hold up to 5,000 properties and be as long as 120,000 characters. A schema over any limit is refused with 400 output_schema_invalid before the run starts, and Sume charges nothing for a refused request (Sume docs: Structured output, read 2026-10-06).

For most video work you will never get near these numbers. A season with a dozen episodes, each a small object, is a few hundred properties. The limits matter when a team tries to put an entire content calendar into one schema.

How the refusal looks

The 400 response carries a details.violations list that names each problem and where it is. Typical violations are an unsupported keyword, an object missing additionalProperties false, a property absent from required, and a $ref that points anywhere but #/$defs/ or SumeMediaFile#. Read the list from top to bottom; one change often clears several entries.

Because the refusal happens before the run, it is the cheapest possible failure. Test a new schema with a throwaway call and a low cap, and fix the violations before you wire it into a bulk queue of a hundred items.

Hard limits on a Sume output schema, read 2026-10-06 against Sume docs
LimitValueResult when exceeded
Nesting depth10400 output_schema_invalid
Properties5,000400 output_schema_invalid
Schema size120,000 characters400 output_schema_invalid
Top-level shapeObject only400 output_schema_invalid

Flatten before you hit a limit

If you are near a limit, the schema is probably doing a job better done outside it. Reuse repeated shapes with $defs and $ref so the schema stays short. Prefer a list of small objects over a deeply nested tree. And separate concerns: one run returns the episode, and your own code assembles the season, rather than one giant schema that tries to hold both.

A quick local check saves a round trip. The script below measures size and depth of a schema before you send it.

import json

def depth(node, d=1):
    if isinstance(node, dict):
        return max([d] + [depth(v, d + 1) for v in node.values()])
    if isinstance(node, list):
        return max([d] + [depth(v, d + 1) for v in node])
    return d

def count_props(node):
    n = 0
    if isinstance(node, dict):
        n += len(node.get("properties", {}))
        n += sum(count_props(v) for v in node.values())
    elif isinstance(node, list):
        n += sum(count_props(v) for v in node)
    return n

schema = json.load(open("schema.json"))
print("chars", len(json.dumps(schema)), "depth", depth(schema),
      "props", count_props(schema))

A caveat on the local check

Treat the script as a rough guard. Sume's own counting is what decides, and your depth may differ by a level or two from how it counts schema keywords versus data nesting. If you are within a few levels of the limit, trust the server and simplify rather than chase the exact boundary.

Bounded lists

Use maxItems on arrays so a runaway list cannot swell the receipt. The same keyword family, minItems, maxItems, minLength, maxLength and pattern, is in the accepted set, and bounding sizes keeps the receipt small enough to stay under the 1 MiB webhook limit.

Test the schema cheaply

The refusal happens before the run starts and Sume charges nothing for it, so the cheapest test of a schema is a call with a low cap and a tiny instruction. If the schema is accepted, a run starts and you can cancel it; if it is refused, you have the list of violations to fix.

Keep schemas in files next to your code, reviewed like any other contract. A schema pasted into a script becomes a schema nobody can find when a field needs to change.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume