Can a Haiku 5.5 agent look at images? Sume's 30-image attachment rules
Anthropic says all current Claude models take image input. Sume accepts up to 30 images, 30 MB each, 500 MB total; Haiku 5.5 is a Format-run pick.

Short answer
Yes on the vendor side. Anthropic's models overview says all current models, Haiku 5.5 included, support text and image input and text output. On Sume's side, the rules for images are the ones on the Agent Completions and Format call pages: a maximum of 30 images, each up to 30 MB, 500 MB for the set, over HTTPS URLs or an asset_id from the workspace. Remember that only Format runs let you choose Haiku 5.5; an Agent Completion runs sume-agent.
The Mistral Large 4 page lists a 1.6B vision encoder, so it also reads images, but Sume does not list it.
Limits table
These are request limits, not model limits, so they do not change with the orchestrator.
| Item | Value | Source |
|---|---|---|
| Images per request | 30 | Agent Completions, Format call |
| Single image | 30 MB max (413 attachment_too_large above it) | Agent Completions errors |
| Whole set | 500 MB max | Agent Completions errors |
| Source | Public HTTPS image_url or workspace asset_id, not both | Agent Completions errors |
| Haiku 5.5 image input | Supported | Anthropic models overview |
| Non-image files | Not available yet | Agent Completions |
Errors that look like model problems
Three failures come from the attachment, not from the model. 400 invalid_attachment means a wrong type, a non-HTTPS URL, both image_url and asset_id, or a source that is not a permitted image. 502 attachment_fetch_failed means Sume could not fetch the image: an unreachable host, hotlink protection, or a non-2xx response. 413 is size. Switching from Astra to Haiku will not fix any of them.
A shot-review use
Attach the frames you want judged and ask for a structured result with output_schema. Sume still parses the final output against your schema, and the images go to the agent. A turn that is only images is allowed; Sume tells the agent to use the attached files when you send no text.
What to put in the image prompt
Image input counts as input tokens, so many attachments add to the prompt length on every turn. Anthropic's pricing page says its vision pricing is set on image size, and a prompt past 100,000 tokens moves Haiku 5.5 to the higher rates. A dozen large frames can matter more than the text of the instruction.
Send the smallest frames that still show what you need, and send each only once: a Format run with previous_run_id continues a conversation, so images already in the thread need not be attached again. State what you want judged, such as 'is the product label legible', in the first line.
One more point on cost: images are priced as input, so the image-heavy turn is usually the first turn, and later turns in the same thread reuse the earlier context. If the cache applies, the repeat cost is the cache read rate and not the full input rate. That is another reason to continue a thread with previous_run_id and not to start a fresh run that re-sends every frame. Check the receipt after the first run to see how many tokens the images actually took, and use that number, not a guess, to plan the next batch.
- Haiku 5.5 over 100k tokens: $0.50 input, $2.50 output per million.
- Keep attachments to the frames that decide the answer.
Sources
Related posts
More in Models
- 1,000 agent turns: Haiku 5.5 about $4, GPT-6 Astra about $400
At list rates, 1,000 turns of 30,000 input and 2,000 output tokens cost about $4 on Haiku 5.5 and $400 on GPT-6 Astra. The arithmetic and the cache effect.
- Higgsfield Soul on Sume: half a cent per image, batches of 1 or 4
higgsfield-soul costs $0.005 at 720p and $0.0075 at 1080p on Sume. It takes batch sizes 1 or 4, no references. Price table and a four-image call.
- Higgsfield Soul on Sume: 1 cent an image, n of 1 or 4, $100 per 10,000
Higgsfield Soul is the cheapest Sume image model at 1 cent. It is text-only, takes n of 1 or 4 and 720p or 1080p, so 10,000 images is 2,500 calls and $100.
- High to max on 100 hero images: GPT Image 2.5 adds $20, xhigh adds $5
Upgrade 100 GPT Image 2.5 hero images from high to xhigh or max: the extra cost on Sume is $5 or $20 at 1024-class size. The arithmetic is shown.
Written by Sume