Thinking effort defaults: Haiku 5.5, Astra and GLM Flash on Sume

Anthropic defaults Haiku 5.5 to medium effort; OpenAI sets no Astra default. Sume's picker offers Astra Low, Medium, High (default High); Haiku, GLM have none.

4 min readSume
All posts

Short answer

Effort defaults differ at three levels. Anthropic's models overview lists Claude Haiku 5.5 with adaptive thinking and a default effort of medium on the Claude API. OpenAI's GPT-6 Astra page lists effort options low, medium, high, xhigh and max and states no default. Inside Sume's agent picker, Astra offers only Low, Medium and High, with High as the default, and Haiku 5.5 and GLM 5.3 Flash expose no effort control at all.

Z.ai's pricing page and the Mistral Large 4 page, read the same day, say nothing about effort defaults, so I make no claim about those two.

Table

The Sume column describes the picker's parameter definitions in the registry, not the raw vendor API.

Effort by model (Sume agent model registry and rate cards on origin/main, read 2026-10-08; vendor pages read 2026-10-08)
ModelVendor defaultVendor optionsSume picker
Claude Haiku 5.5medium (Claude API)Set explicitly to change; adaptive thinkingNo effort parameter on the row
GPT-6 AstraNot stated on the model pagelow, medium, high, xhigh, maxLow, Medium, High; default High; Fast checkbox off by default
GLM 5.3 FlashNot stated on the Z.ai pricing pageNot statedNo effort parameter on the row
Mistral Large 4 previewNot statedNot statedNot in Sume

Why the picker stops at High

The registry comment says xhigh, max, minimal and none are wire spellings that environment settings accept so that a bad value cannot take a turn down, and that max in particular spends the whole pre-text window reasoning. They are not product choices, so a per-turn chip does not offer them. That is a design call by Sume, not an OpenAI limit.

The Fast checkbox is a separate axis: OpenAI lists Fast at 2x Standard, and Sume's card mirrors that.

What to do for a video agent

If cost matters, the default you inherit is the one that pays: Astra at High effort bills reasoning as output tokens at $50 per million. Pick Low when the task is a lookup or a reformat, and keep High for planning a multi-scene edit. For Haiku 5.5 there is nothing to set on Sume; the vendor default applies unless Sume's runtime passes something else, which the registry does not show.

Reading a bill for effort

Effort shows up on the bill as output tokens, not as a line item. On Astra, a High-effort turn that writes 6,000 tokens of reasoning and answer costs 6,000 x 50 / 1M = $0.30 in output, while a Low-effort turn that writes 1,500 costs $0.075. The same input of 30,000 tokens adds $0.30 either way. Those token counts are illustrative; Sume's receipt is the place to see real ones.

Because Sume's picker default is High for Astra, an unattended run that never touches the setting pays for the deepest of the three levels it offers. Haiku 5.5 at medium output pricing of $0.50 per million makes the same extra 4,500 tokens cost $0.00225.

Where does that leave the choice? For a planning turn that routes between tools, a lower level is usually enough, and you can set it in the picker for a thread. For a turn that has to reconcile several conflicting constraints, a higher level may pay for itself. Sume's picker does not decide for you; it exposes the levels and a default. Check the effort on the thread before a long unattended run, and set it deliberately.

  • Astra extra 4,500 output tokens: 4,500 x 50 / 1M = $0.225.
  • Haiku 5.5 extra 4,500 output tokens: 4,500 x 0.50 / 1M = $0.00225.

Sources

Related posts

More in Models

All Models posts

Written by Sume