Claude structured output refusal or max_tokens: what Sume does

Claude can return a 200 that does not match your schema on refusal or max_tokens. Sume reports a failed projection as output_error on the receipt instead.

5 min readSume
All posts

A Claude structured output is not guaranteed to match your schema when the response stops with stop_reason: "refusal" or stop_reason: "max_tokens". Claude's docs say a refusal still returns HTTP 200 and still bills tokens, and a max_tokens stop leaves incomplete JSON that will not validate. Sume has neither branch: a Format run binds your schema to the finished run, and the failure you handle is output_error on the receipt.

This page compares the two failure models so you know which check to write when you move a schema from output_config.format to Sume's output_schema. Claude facts are from Claude's structured outputs page; Sume facts are from Structured output and Errors and spend.

What does Claude return when it refuses or runs out of tokens?

On Claude, you pass output_config.format with type: "json_schema" and read the model's text as JSON. The page names two stop reasons where that guarantee does not hold, and both come back as a normal 200 response.

Claude stop reasons that break the schema guarantee (read 2026-10-02)
stop_reasonSchema matchBillingWhat to do
refusalOutput may not matchTokens still billedRead the refusal message, do not parse as your type
max_tokensOutput is cut off and will not validateBilled as usualRetry with a higher max_tokens
end_turnMatches the schemaBilled as usualParse it

What replaces those branches on a Sume Format run?

Sume does not constrain a model's reply. Your schema is handed to the run as a tool it must call before finishing, and if the run does not submit a valid object, a separate constrained pass builds one from the run's media and closing text. The receipt tells you which happened in filled_by: agent or projection.

That is why the OpenAI-to-Sume mapping in the docs says a refusal on the message becomes output_error on the receipt, and why incomplete_details.reason: "max_output_tokens" is marked not applicable: the projection is small and bounded, so there is no truncated-JSON case. The same applies when you port from Claude. There is no stop_reason to switch on, and no partial JSON string to repair.

Do not read that as "cannot fail". A run can finish with status: "failed" and output: null because the projection was not satisfied.

Which Sume field do I check instead?

Check output_error before reading output, as the structured output page tells you to. Each code says why no object was accepted, and details.harvested counts the media the run actually made, by type.

These are the codes the docs list for the projection and its gates:

  • output_schema_unsatisfied: the projection did not match your schema, or referenced media the run did not produce.
  • output_extraction_failed: the projection could not run; with reason: harvest_unavailable the run stays completed and the receipt fills in on your next read.
  • unattended_blocked, deliverable_missing, primary_output_missing, agent_reported_failure: the run stopped, never made the declared media, left your primary_output_key empty, or reported that it did not deliver.
  • The docs call the set open, so branch on the codes you handle and fall through on the rest.

What does a safe reader look like?

This reader refuses to treat a receipt as delivered unless the run is not failed, output_error is empty and primary_output_url is set. It reads the run receipt documented in Runs and results.

import os, sys, requests

run_id = sys.argv[1]
r = requests.get(
    f"https://api.sume.com/v1/format-runs/{run_id}",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    timeout=30,
)
r.raise_for_status()
run = r.json()["data"]
err = run.get("output_error")
if run["status"] == "failed" or err:
    code = (err or run.get("error") or {}).get("code")
    print("not delivered:", code)
elif run.get("primary_output_url"):
    print("delivered:", run["primary_output_url"], run.get("filled_by"))
else:
    print("still running or no deliverable:", run["status"])

What does Sume not do that Claude does?

Sume does not stream partial JSON; output appears once, on the terminal receipt. It has no JSON mode without a schema either: you bind a schema or take the Format's built-in one. And because the schema constrains how a finished run is read back, it cannot make a Format produce a video it never made. If you need Claude to write the JSON itself, call Claude directly; if you need the deliverable media plus a typed receipt, bind output_schema and handle output_error.

For the failure codes in full, see Format run failure codes.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume