Sume's caption aligner runs on Haiku 5.5: what it means for you
A sidecar LLM step in Sume's video captions moved from Haiku 4.5 to Haiku 5.5. The card is $0.10 in, $0.50 out per million; why the change is small.

Short answer
Besides the agent picker, Sume uses Claude Haiku 5.5 inside a product feature. The repo's sidecar rate cards describe the video caption aligner as running on Haiku 5.5 since #10454, replacing Haiku 4.5. The aligner is a small language-model step; it is not something you select, and the docs for standalone video captions do not describe it, so there is nothing new to configure.
Anthropic's models overview describes Haiku 5.5 as for "high-volume, latency-sensitive tasks such as classification, extraction, and routing", which matches a small alignment job better than a big planning one.
The card
Sume prices the aligner's model call on the Haiku 5.5 base tier.
| Item | Haiku 5.5 | Haiku 4.5 |
|---|---|---|
| Input | $0.10 | $1.00 |
| Output | $0.50 | $5.00 |
| Cache read | $0.01 | $0.10 |
| Cache write | $0.125 | $1.25 |
Size of the effect
A hypothetical alignment call with a 4,000-token prompt and a 500-token answer costs 4,000 x 0.10 / 1M = $0.0004 plus 500 x 0.50 / 1M = $0.00025, so $0.00065 on Haiku 5.5. On 4.5 the same call would be $0.004 plus $0.0025, so $0.0065. Both are fractions of a cent per call; I do not know the real token counts of the aligner, so treat these as sizing, not as a measured cost.
Captions themselves are billed under their own price in the catalog. Nothing in the public docs says that the aligner's model tokens appear as a separate line.
What to do
Nothing, unless you compare caption output quality across weeks. A model swap can change how scripts line up with speech. If a caption alignment result looks different from last week's, the model change in #10454 is one thing to rule out before blaming the audio.
Where else Haiku shows up
Haiku 5.5 appears in Sume's code in one more place: it is a catalog row in the agent picker, enabled, labeled Haiku 5.5, with the tooltip 'Anthropic's fastest model, for high-volume, latency-sensitive work'. The two uses share a rate card, so a price change by Anthropic would move both.
Anthropic's 100,000-token step does not matter for the aligner, since a caption alignment prompt is far below it. It matters for the agent row, where a long session can cross it.
The practical lesson is to look at where a model swap touches you before you celebrate or worry about it. A caption alignment call sends a transcript and gets timing back, so its input is small and its output is short. At the Haiku 5.5 rates, even a thousand such calls cost a few cents. The swap from 4.5 to 5.5 is therefore a quality and latency question for captions, not a billing one, and the billing impact only shows when you pick Haiku 5.5 as the main agent model for long sessions.
- Picker row: Haiku 5.5, not the recommended pick.
- Aligner: sidecar rate card keyed claude-haiku-5.5.
Sources
Related posts
More in Models
- Sume's GLM 5.3 Flash card is $0.075 in; Z.ai's page says $0.15
Sume's registry prices GLM 5.3 Flash at $0.075 input and $0.25 output per million. Z.ai's pricing page lists $0.15 and $0.50. Here is what to do with that gap.
- Sume's image router lists 17 model ids on Oct 8, 2026; no Hy Image
The public Sume catalog on 2026-10-08 lists 17 model ids under Image Router. None is Tencent's Hy Image 3.5 Preview. See the list and the check.
- sume/music-auto or pinned lyria-3.5: what changes in the job
Music Router takes sume/music-auto (Lyria 3.5 today), lyria-3.5 or lyria-3-pro. What job.model and routed_model echo back, and why the price stays the same.
- Does the Sume spend cap cover the agent's own model tokens?
No: Sume's cap excludes the agent's LLM turn; debited_usd_micros includes it. At list rates that turn is about $0.004 on Haiku 5.5 and $0.40 on Astra.
Written by Sume