GPT-6.1 Sol structured outputs or a Sume Format output_schema?
GPT-6.1 Sol supports structured outputs and function calling in OpenAI's API. A Sume Format run adds media, a spend cap and a receipt. Pick by the job.

Choose the model's structured outputs when the answer is text that fits in one request. Choose a Sume Format run when the job makes media or calls tools over several steps and you want a typed result plus a spend ceiling. OpenAI's GPT-6.1 Sol page lists structured outputs and function calling as supported features, so a plain extraction needs no run at all.
The unit of work
The difference is the unit of work. A model call returns tokens. A Format run is one agent turn in a sandbox: it can generate files, and the receipt lists artifacts, usage and output.
| Need | Model structured outputs | Sume Format run |
|---|---|---|
| Typed text answer | Fits | Works, but heavier |
| Generated video or image | Not the job | artifacts[] and primary_output_url |
| Spend ceiling | Your own code | generation_spend_cap_usd on the request |
| Retry safety | Your own code | Idempotency-Key header |
| Result shape | Schema you send | output_schema, filled_by, output_error |
What a run adds
A run is asynchronous. It answers 202 with a receipt, and you read status_url or take one terminal webhook. See Calling a Format for the request fields.
Check the docs before you ship
Sume's limits and field names change faster than blog posts do. Read the linked docs pages for the current request fields before you ship, and send a dry_run or a low spend cap on your first real call.
Sources
Related posts
More in Comparisons
- GPT Image 2.5 mask edit, then an Ideogram 4.5 text pass: one chain
Use GPT Image 2.5 for the masked region change and Ideogram 4.5 for the text pass, both on POST /v1/images. When it beats one model, and what each step bills.
- GPT Image 2.5 medium is 6x cheaper than Nano Banana 2 1K on Sume
On Sume, GPT Image 2.5 at medium costs $0.0165 for 1024x1024 and Nano Banana 2 1K costs $0.10, a 6.1x gap. Ratios at 2K and 4K, and what the price omits.
- Griffin clones a voice from about 10 seconds: what Sume does for voice
Tavus says Griffin can clone a voice from about 10 seconds of audio. Sume's TTS docs describe choosing a voice, not training one. Plan around that gap.
- Griffin-Lite or a rendered avatar clip for Q4? A decision table
Tavus Griffin-Lite is a research preview for invited testers. Sume Avatar 1.0 renders scripted clips from $11.04 per minute. Pick by the job, not the demo.
Written by Sume