output_schema max_string_length: descriptions can spend 120,000

A Sume Format output_schema has a document-wide 120,000-character string budget, plus depth 10, 5,000 properties and 1,000 values per enum.

4 min readSume
All posts

A Sume output_schema is rejected with max_string_length when the strings in the whole document add up to more than 120,000 characters. The budget is summed over every property name, key and string value in the schema, so a large schema full of long description annotations can exhaust it even when no single string looks big. The three other size limits are nesting depth 10, 5,000 properties in total, and 1,000 values per enum.

What are the schema size limits?

These four limits apply at submit. A schema over any of them fails with 400 output_schema_invalid and a details.violations[] entry, and nothing runs, so nothing is charged.

The Scheduled create page lists the first three limits (10 levels, 5000 properties, 1000 enum values) for schemas bound in the dashboard. The string budget is documented on the structured output page.

output_schema limits, from docs.sume.com/formats/structured-output (read 2026-10-03)
LimitValueViolation rule
Nesting depth10 levelsmax_depth
Total properties5000, counted across the whole documentmax_properties
Enum values1000 per enummax_enum_values
Total string length120,000 characters, summed over every property name, key and string valuemax_string_length

Why does it count the whole document?

The last limit is a document-wide budget rather than a per-field cap. That is what surprises people who generate a schema from a long spec: the descriptions are annotations to you, but they are strings in the document, and they count toward the total.

Depth works differently from what you might expect. It counts literal nesting in the document, so a $defs entry that references itself does not consume it. If your schema is deep because it models a tree, a self-referencing definition is the cheaper shape.

How do I find out how close I am?

Count before you submit. This rough local check adds the length of every key and string value in a schema file. It is an estimate for planning, not the platform's exact counter, so leave headroom.

import json, sys

def total(node):
    if isinstance(node, dict):
        return sum(len(k) + total(v) for k, v in node.items())
    if isinstance(node, list):
        return sum(total(v) for v in node)
    if isinstance(node, str):
        return len(node)
    return 0

schema = json.load(open(sys.argv[1]))
print(total(schema), "of 120000")

What should I do when the budget runs out?

Trim prose first. Long description and examples text is the usual cause, and guidance for the agent belongs in the Format's SKILL.md or in the run instruction, not in the schema. Then look at enums: a 1,000-value cap per enum is generous, but a lookup table pasted as an enum also spends string budget.

Read the violation rather than guessing. Every entry is { path, rule, message }; rule is a stable token you can switch on, and every problem is reported in one 400, so one rejected submit is enough to fix the schema. Two related facts: strict: false does not relax any of these limits, and there is no JSON-mode escape hatch, so the only options are a schema inside the subset or the built-in one.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume