Can a Haiku 5.5 agent look at images? Sume's 30-image attachment rules

Anthropic says all current Claude models take image input. Sume accepts up to 30 images, 30 MB each, 500 MB total; Haiku 5.5 is a Format-run pick.

4 min readSume
All posts

Short answer

Yes on the vendor side. Anthropic's models overview says all current models, Haiku 5.5 included, support text and image input and text output. On Sume's side, the rules for images are the ones on the Agent Completions and Format call pages: a maximum of 30 images, each up to 30 MB, 500 MB for the set, over HTTPS URLs or an asset_id from the workspace. Remember that only Format runs let you choose Haiku 5.5; an Agent Completion runs sume-agent.

The Mistral Large 4 page lists a 1.6B vision encoder, so it also reads images, but Sume does not list it.

Limits table

These are request limits, not model limits, so they do not change with the orchestrator.

Attachment limits (Sume docs and vendor pages, read 2026-10-08)
ItemValueSource
Images per request30Agent Completions, Format call
Single image30 MB max (413 attachment_too_large above it)Agent Completions errors
Whole set500 MB maxAgent Completions errors
SourcePublic HTTPS image_url or workspace asset_id, not bothAgent Completions errors
Haiku 5.5 image inputSupportedAnthropic models overview
Non-image filesNot available yetAgent Completions

Errors that look like model problems

Three failures come from the attachment, not from the model. 400 invalid_attachment means a wrong type, a non-HTTPS URL, both image_url and asset_id, or a source that is not a permitted image. 502 attachment_fetch_failed means Sume could not fetch the image: an unreachable host, hotlink protection, or a non-2xx response. 413 is size. Switching from Astra to Haiku will not fix any of them.

A shot-review use

Attach the frames you want judged and ask for a structured result with output_schema. Sume still parses the final output against your schema, and the images go to the agent. A turn that is only images is allowed; Sume tells the agent to use the attached files when you send no text.

What to put in the image prompt

Image input counts as input tokens, so many attachments add to the prompt length on every turn. Anthropic's pricing page says its vision pricing is set on image size, and a prompt past 100,000 tokens moves Haiku 5.5 to the higher rates. A dozen large frames can matter more than the text of the instruction.

Send the smallest frames that still show what you need, and send each only once: a Format run with previous_run_id continues a conversation, so images already in the thread need not be attached again. State what you want judged, such as 'is the product label legible', in the first line.

One more point on cost: images are priced as input, so the image-heavy turn is usually the first turn, and later turns in the same thread reuse the earlier context. If the cache applies, the repeat cost is the cache read rate and not the full input rate. That is another reason to continue a thread with previous_run_id and not to start a fresh run that re-sends every frame. Check the receipt after the first run to see how many tokens the images actually took, and use that number, not a guess, to plan the next batch.

  • Haiku 5.5 over 100k tokens: $0.50 input, $2.50 output per million.
  • Keep attachments to the frames that decide the answer.

Sources

Related posts

More in Models

All Models posts

Written by Sume