The built-in Format output schema: sume/action-run-output/v1

No output_schema on a Format run returns sume/action-run-output/v1: text plus four media arrays, filled without a model, so it cannot fail like a custom schema.

4 min readSume
All posts

If you bind no output_schema, a Format run returns output in the built-in sume/action-run-output/v1 shape: a nullable text plus four always-present arrays, images, videos, audio and files. It is filled deterministically from the run's generated media and final text, with no model involved, so it cannot fail the way a custom schema can.

This is from Structured output, read on 2026-10-02.

What does the output look like?

Each media entry is a SumeMediaFile-shaped object, as in the schema reference.

{
  "text": "Short summary written by the Agent.",
  "images": [],
  "videos": [
    {
      "type": "video",
      "url": "https://media.sume.com/...",
      "content_type": "video/mp4",
      "file_name": "teaser.mp4",
      "size_bytes": 4210233,
      "width": 1080,
      "height": 1920,
      "duration_ms": 12000,
      "expires_at": null
    }
  ],
  "audio": [],
  "files": []
}

How do I know which schema applied?

The receipt's output_schema.source says.

Read 2026-10-02 from docs.sume.com
sourceMeaning
defaultNothing was bound; output is the built-in schema
action_defaultThe Format's own schema bound in the dashboard
request_overrideThe output_schema sent on this run request

When should I just use it?

When you want the media and do not need a custom shape. The docs say it is already enough to ship on. A per-request output_schema overrides a Format's bound default for that run, and the built-in is what you fall back to when neither exists.

The primary_output_url for the built-in schema falls back through videos, then images, audio, files, using the first non-empty array.

What are the limits of it?

It gives you files and a paragraph, not your own fields. Ids, titles and labels from your records are not in it, and input never reaches the output, so keep your identifiers on your side keyed by data.id or your Idempotency-Key.

If you need typed fields such as a headline and alt text next to the image, bind a custom schema, accept that it can return output: null with output_error, and read artifacts[] as the fallback. The built-in does not need that handling for shape, though an empty array still means the run made none of that type.

Is it the same on a failed run?

A failed run still publishes what it generated, and primary_output_url is null on every non-completed run. Read status and error first, then the arrays.

Tips for reading it

Switching to a custom schema later is a per-request change: send output_schema and the receipt reports source: request_override for that run only.

  • Always check each array's length before indexing; the four arrays are present but may be empty.
  • Treat text as nullable.
  • Use videos[0].url for a single-clip Format, or set primary_output_key to videos semantics via the fallback order.
  • Persist artifacts[] too, since it lists every durable file the run generated, with checksums.

How is it different from a custom schema?

A custom schema is filled by the run itself when it submits a valid object (filled_by: agent), or by a separate constrained pass at temperature 0 from the run's media and the first 8000 characters of its closing text (filled_by: projection). Both are gated for URLs and durations.

The built-in schema is filled without a model, directly from the generated media and final text. That is why the docs say it cannot fail the way a custom schema can. The price is flexibility, since the shape is fixed.

Does it carry my own ids?

No. The projection never sees your input, and the built-in schema has no field for your identifiers. Key your records by the run's data.id or the Idempotency-Key you sent, and write the media URLs against that record when the terminal webhook or a poll arrives.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume