Agent Completions input_image: send a photo in messages[]

Send an image to Sume's Agent Completions as an input_image content part inside messages[], merge it with attachments, and get a caption back as typed output.

5 min readSume
All posts

To show Sume's agent a picture through Agent Completions, put an {"type": "input_image", "image_url": "https://..."} part in a user message's content array, next to an input_text part if you have instructions. The same item shape also works as a top-level attachments array, and Sume merges both sources into one list. Images are the only attachment type today.

Below: the request, the limits, the errors, and how attachments combine with output_schema.

What does the request look like?

From the Agent Completions docs: content accepts a string or an OpenAI-style array of parts, and also accepts input_text as an alias for text. The example below asks for a caption and binds a schema so output comes back typed. It must include generation_spend_cap_usd, which has no default.

curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": [
      {"type": "input_text", "text": "Describe this product shot."},
      {"type": "input_image", "image_url": "https://cdn.example.com/shot.jpg"}
    ]}],
    "output_schema": {"name": "caption", "schema": {
      "type": "object", "required": ["caption"],
      "additionalProperties": false,
      "properties": {"caption": {"type": "string"}}}},
    "generation_spend_cap_usd": 2
  }'

What are the attachment rules and limits?

An image-only turn is fine: leave out the text part and the agent is told to use the attached files. Item shape, limits, upload path and error codes are identical to Format runs, as the docs say, and the table collects the ones that matter.

Agent Completions attachment rules (read 2026-10-02)
RuleValue
Images per requestUp to 30
Size per image30 MB maximum
Total size500 MB maximum
Sourceimage_url over HTTPS, or an asset_id from your workspace, never both
Attachment typeinput_image only; PDFs and other files are not available yet

Which errors point at the image?

Three codes are specific to attachments. 400 invalid_attachment covers a wrong type, a missing or non-HTTPS URL, both image_url and asset_id on one item, or a source that is not an allowed image. 400 attachment_not_found means the asset_id is unknown in this workspace. 413 attachment_too_large means an image over 30 MB or a set over 500 MB. A fourth, 502 attachment_fetch_failed, means Sume could not fetch the URL: an unreachable host, hotlink protection, or a non-2xx answer. For that one, upload the image first and send its asset_id, or host it somewhere that serves it without checking a referrer.

How do images and output_schema work together?

They compose: the images reach the agent, and the run's output is parsed against your schema after the run completes. That is the shape for caption, alt text or tagging jobs where a database column wants a string and not a paragraph. The accepted call returns 202 with a run receipt, not a choices[] answer, and you poll status_url until the run reaches a terminal status. The create call is asynchronous because an agent turn may open a sandbox, call tools and generate media.

Two limits to plan around. Every completion runs in a fresh thread, and the docs list continuing a prior thread with thread_id as not available yet. And streaming and a synchronous OpenAI-compatible response are not available, so a plug-in that expects choices[] will not work unchanged.

URL or asset_id: which should you send?

Use image_url when the file is already public over HTTPS. The Format docs say Sume fetches it at create time, so it must be reachable without authentication, and it checks the real type and size and copies the image into durable storage. A broken or private image therefore fails the create with an error you can act on, instead of killing the run minutes later. The accepted types are JPEG, PNG, WebP, GIF and AVIF.

Use asset_id when the image sits behind your own auth or a hotlink rule. Upload it through the Assets API first, once it is ready in the same workspace, and send the id; an asset_id or a URL already on media.sume.com is not copied again. Send one or the other on an item, never both. An optional filename sets the label the agent sees, which defaults to the URL's basename; name files descriptively when you attach several, so the instruction can refer to packshot.png rather than "the second image".

Last, remember the Agent Completions spend cap. An image-only caption job rarely generates media, but the cap is still required on every request, so send generation_spend_cap_usd even for a read-only task.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume