Thinking effort defaults: Haiku 5.5, Astra and GLM Flash on Sume
Anthropic defaults Haiku 5.5 to medium effort; OpenAI sets no Astra default. Sume's picker offers Astra Low, Medium, High (default High); Haiku, GLM have none.

Short answer
Effort defaults differ at three levels. Anthropic's models overview lists Claude Haiku 5.5 with adaptive thinking and a default effort of medium on the Claude API. OpenAI's GPT-6 Astra page lists effort options low, medium, high, xhigh and max and states no default. Inside Sume's agent picker, Astra offers only Low, Medium and High, with High as the default, and Haiku 5.5 and GLM 5.3 Flash expose no effort control at all.
Z.ai's pricing page and the Mistral Large 4 page, read the same day, say nothing about effort defaults, so I make no claim about those two.
Table
The Sume column describes the picker's parameter definitions in the registry, not the raw vendor API.
| Model | Vendor default | Vendor options | Sume picker |
|---|---|---|---|
| Claude Haiku 5.5 | medium (Claude API) | Set explicitly to change; adaptive thinking | No effort parameter on the row |
| GPT-6 Astra | Not stated on the model page | low, medium, high, xhigh, max | Low, Medium, High; default High; Fast checkbox off by default |
| GLM 5.3 Flash | Not stated on the Z.ai pricing page | Not stated | No effort parameter on the row |
| Mistral Large 4 preview | Not stated | Not stated | Not in Sume |
Why the picker stops at High
The registry comment says xhigh, max, minimal and none are wire spellings that environment settings accept so that a bad value cannot take a turn down, and that max in particular spends the whole pre-text window reasoning. They are not product choices, so a per-turn chip does not offer them. That is a design call by Sume, not an OpenAI limit.
The Fast checkbox is a separate axis: OpenAI lists Fast at 2x Standard, and Sume's card mirrors that.
What to do for a video agent
If cost matters, the default you inherit is the one that pays: Astra at High effort bills reasoning as output tokens at $50 per million. Pick Low when the task is a lookup or a reformat, and keep High for planning a multi-scene edit. For Haiku 5.5 there is nothing to set on Sume; the vendor default applies unless Sume's runtime passes something else, which the registry does not show.
Reading a bill for effort
Effort shows up on the bill as output tokens, not as a line item. On Astra, a High-effort turn that writes 6,000 tokens of reasoning and answer costs 6,000 x 50 / 1M = $0.30 in output, while a Low-effort turn that writes 1,500 costs $0.075. The same input of 30,000 tokens adds $0.30 either way. Those token counts are illustrative; Sume's receipt is the place to see real ones.
Because Sume's picker default is High for Astra, an unattended run that never touches the setting pays for the deepest of the three levels it offers. Haiku 5.5 at medium output pricing of $0.50 per million makes the same extra 4,500 tokens cost $0.00225.
Where does that leave the choice? For a planning turn that routes between tools, a lower level is usually enough, and you can set it in the picker for a thread. For a turn that has to reconcile several conflicting constraints, a higher level may pay for itself. Sume's picker does not decide for you; it exposes the levels and a default. Check the effort on the thread before a long unattended run, and set it deliberately.
- Astra extra 4,500 output tokens: 4,500 x 50 / 1M = $0.225.
- Haiku 5.5 extra 4,500 output tokens: 4,500 x 0.50 / 1M = $0.00225.
Sources
Related posts
More in Models
- TTS Router ids: sonic-3.6 to sonic-preview, one price book
Sume TTS Router lists sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview, all at the TTS 1.0 character price. What each id is and when to pin it.
- 12 Sume video ids: which need a source video, which take text
Nine of Sume's twelve video ids take a text prompt. grok-imagine-video-1.5 needs an image; higgsfield-genjutsu and h3-max-recast need a video. Ranges per id.
- A 20-turn agent session: Haiku 5.5 $0.03, Astra $3.10 with caching
Cache write on turn one, reads after: a 20-turn session with a 35k-token cached prefix costs about $0.031 on Haiku 5.5 and $3.10 on GPT-6 Astra. Arithmetic.
- US company and open video weights: H3 is out, three others differ
MiniMax H3 weights exclude the US by license. HunyuanVideo, Wan 2.2 and LTX-2.5 read differently. A territory table and a hosted route on Sume.
Written by Sume