Agent Completion or Format run: where can you pick Haiku 5.5?
Only a Format run takes a catalog model id. An Agent Completion accepts model sume-agent and nothing else, so Haiku 5.5 is a Format-run choice.

The difference in one line
If you want to choose Claude Haiku 5.5 as the orchestrator of a Sume run, use a Format run. If you call POST /v1/agent/completions, the only value model accepts is sume-agent, and any other value is a 400 invalid_request. Omitting it gives you the same agent.
This matters this week because Haiku 5.5, GLM 5.3 and Mistral Large 4 all arrived within two days, and people will try to paste their ids into the nearest endpoint.
Side by side
Both surfaces run the same agent and return the same receipt shape, per the Agent Completions page. The difference is where the instruction lives and which fields exist.
| Question | Agent Completion | Format run |
|---|---|---|
| Endpoint | POST /v1/agent/completions | POST /v1/formats/{handle}/{slug}/runs |
| model field | Only sume-agent; other values give 400 invalid_request | Agents catalog id; id outside the catalog gives 400 invalid_request |
| Saved workflow | None; you send the task each call | The Format stores how to do it |
| Spend cap | generation_spend_cap_usd required, no default | Optional; inherits the Format's cap; max $500 |
| Streaming / choices[] response | Not available | Not applicable; receipt and polling |
What to do with a Haiku idea
If the task changes on every call and you cannot save a Format, you cannot select Haiku 5.5 through Agent Completions. You can still save a small Format whose instructions take an input object, then run it with model set. Whether that fits depends on whether the task really changes shape or only its data.
If your cost concern is the orchestrator's tokens, compare it with the spend cap first: the Format docs say usage.billable_amount_usd_micros does not include the agent's own LLM turn, while usage.debited_usd_micros does.
Error you will see
Sending model: "claude-haiku-5-5" to Agent Completions returns 400 invalid_request; the error table on the Agent Completions page lists a model other than sume-agent as a cause. Nothing runs and nothing is charged for a rejected create.
Making the choice
Pick by who knows the task ahead of time. If the task is a repeatable recipe, such as a vertical teaser from a product brief where only the product and the copy change, a Format is the natural home: the Format stores how, the run supplies input, and you can set model. If every call is a one-off prompt and you only care that an agent does it, Agent Completions is the lighter path, and you accept that Sume chooses the orchestrator.
There is a cost angle to the choice as well. Haiku 5.5 lists at $0.10 input and $0.50 output per million tokens on Anthropic's pricing page, read 2026-10-08. At 30,000 input and 2,000 output tokens, that turn is 30,000 x 0.10 / 1M = $0.003 plus 2,000 x 0.50 / 1M = $0.001, or $0.004. Because the media usually costs far more than the turn, most teams will pick the surface for its workflow fit and not for the orchestrator price.
- Choose Agent Completions for ad-hoc prompts with a required spend cap and no saved state.
- Choose a Format run when you need a model id, a stored recipe,
previous_run_idcontinuation or bulk runs.
Sources
Related posts
More in Models
- AI video news Oct 1-8, 2026: Vidu Q4 Preview, Kling 4.0, Sume coverage
Vidu Q4 Preview launched Oct 7 and kling.ai headlines Kling 4.0. We checked Runway, Luma and Veo, and which of them Sume lists.
- Audio API launches, Oct 1-8 2026: dates, prices and what Sume covers
A dated calendar of speech, voice and music releases from Sep 28 to Oct 8, 2026, each with the number the vendor published, and which ones Sume lists.
- background: transparent on the Sume image API: two GPT rows only
Transparent PNGs from the Sume image API need background: transparent, which only the two GPT Image 2.5 rows list. See the request, price and the 400 elsewhere.
- Best image-to-video model, October 2026: Wan 3.0 is 7th at 1,164 Elo
On the AA image-to-video board MiniMax H3 Max leads at 1,195 and Wan 3.0 is seventh at 1,164. Which of the top 12 you can call on Sume, with per-minute prices.
Written by Sume