Claude structured output refusal or max_tokens: what Sume does
Claude can return a 200 that does not match your schema on refusal or max_tokens. Sume reports a failed projection as output_error on the receipt instead.

A Claude structured output is not guaranteed to match your schema when the response stops with stop_reason: "refusal" or stop_reason: "max_tokens". Claude's docs say a refusal still returns HTTP 200 and still bills tokens, and a max_tokens stop leaves incomplete JSON that will not validate. Sume has neither branch: a Format run binds your schema to the finished run, and the failure you handle is output_error on the receipt.
This page compares the two failure models so you know which check to write when you move a schema from output_config.format to Sume's output_schema. Claude facts are from Claude's structured outputs page; Sume facts are from Structured output and Errors and spend.
What does Claude return when it refuses or runs out of tokens?
On Claude, you pass output_config.format with type: "json_schema" and read the model's text as JSON. The page names two stop reasons where that guarantee does not hold, and both come back as a normal 200 response.
| stop_reason | Schema match | Billing | What to do |
|---|---|---|---|
| refusal | Output may not match | Tokens still billed | Read the refusal message, do not parse as your type |
| max_tokens | Output is cut off and will not validate | Billed as usual | Retry with a higher max_tokens |
| end_turn | Matches the schema | Billed as usual | Parse it |
What replaces those branches on a Sume Format run?
Sume does not constrain a model's reply. Your schema is handed to the run as a tool it must call before finishing, and if the run does not submit a valid object, a separate constrained pass builds one from the run's media and closing text. The receipt tells you which happened in filled_by: agent or projection.
That is why the OpenAI-to-Sume mapping in the docs says a refusal on the message becomes output_error on the receipt, and why incomplete_details.reason: "max_output_tokens" is marked not applicable: the projection is small and bounded, so there is no truncated-JSON case. The same applies when you port from Claude. There is no stop_reason to switch on, and no partial JSON string to repair.
Do not read that as "cannot fail". A run can finish with status: "failed" and output: null because the projection was not satisfied.
Which Sume field do I check instead?
Check output_error before reading output, as the structured output page tells you to. Each code says why no object was accepted, and details.harvested counts the media the run actually made, by type.
These are the codes the docs list for the projection and its gates:
output_schema_unsatisfied: the projection did not match your schema, or referenced media the run did not produce.output_extraction_failed: the projection could not run; withreason: harvest_unavailablethe run stayscompletedand the receipt fills in on your next read.unattended_blocked,deliverable_missing,primary_output_missing,agent_reported_failure: the run stopped, never made the declared media, left yourprimary_output_keyempty, or reported that it did not deliver.- The docs call the set open, so branch on the codes you handle and fall through on the rest.
What does a safe reader look like?
This reader refuses to treat a receipt as delivered unless the run is not failed, output_error is empty and primary_output_url is set. It reads the run receipt documented in Runs and results.
import os, sys, requests
run_id = sys.argv[1]
r = requests.get(
f"https://api.sume.com/v1/format-runs/{run_id}",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
run = r.json()["data"]
err = run.get("output_error")
if run["status"] == "failed" or err:
code = (err or run.get("error") or {}).get("code")
print("not delivered:", code)
elif run.get("primary_output_url"):
print("delivered:", run["primary_output_url"], run.get("filled_by"))
else:
print("still running or no deliverable:", run["status"])What does Sume not do that Claude does?
Sume does not stream partial JSON; output appears once, on the terminal receipt. It has no JSON mode without a schema either: you bind a schema or take the Format's built-in one. And because the schema constrains how a finished run is read back, it cannot make a Format produce a video it never made. If you need Claude to write the JSON itself, call Claude directly; if you need the deliverable media plus a typed receipt, bind output_schema and handle output_error.
For the failure codes in full, see Format run failure codes.
Sources
Related posts
More in Formats
- Sume Format API staging: api.dev.sume.com and the 401 on a wrong host
Sume serves the Format API on api.sume.com and api.dev.sume.com with the same routes. A key works only on its own host; the other answers 401 unauthorized.
- Format bulk queue: no queue webhook, no cancel-queue endpoint
A Sume bulk queue has no webhook and no cancel call. Poll the queue, put webhooks on items, and cancel the child runs one by one with their run ids.
- Format Contents API: read the whole package, commit many files at once
Read a Sume Format package with ?recursive=1 and write several files as one commit and one version bump. A change set, not the package; deletes stay separate.
- Why your order_id comes back null in a Sume structured output
If a Sume run's filled_by is projection, the fallback never sees your input or instruction, so an order_id you sent comes back null. Keep ids on your side.
Written by Sume