Agent Completions input_image: send a photo in messages[]
Send an image to Sume's Agent Completions as an input_image content part inside messages[], merge it with attachments, and get a caption back as typed output.

To show Sume's agent a picture through Agent Completions, put an {"type": "input_image", "image_url": "https://..."} part in a user message's content array, next to an input_text part if you have instructions. The same item shape also works as a top-level attachments array, and Sume merges both sources into one list. Images are the only attachment type today.
Below: the request, the limits, the errors, and how attachments combine with output_schema.
What does the request look like?
From the Agent Completions docs: content accepts a string or an OpenAI-style array of parts, and also accepts input_text as an alias for text. The example below asks for a caption and binds a schema so output comes back typed. It must include generation_spend_cap_usd, which has no default.
curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": [
{"type": "input_text", "text": "Describe this product shot."},
{"type": "input_image", "image_url": "https://cdn.example.com/shot.jpg"}
]}],
"output_schema": {"name": "caption", "schema": {
"type": "object", "required": ["caption"],
"additionalProperties": false,
"properties": {"caption": {"type": "string"}}}},
"generation_spend_cap_usd": 2
}'What are the attachment rules and limits?
An image-only turn is fine: leave out the text part and the agent is told to use the attached files. Item shape, limits, upload path and error codes are identical to Format runs, as the docs say, and the table collects the ones that matter.
| Rule | Value |
|---|---|
| Images per request | Up to 30 |
| Size per image | 30 MB maximum |
| Total size | 500 MB maximum |
| Source | image_url over HTTPS, or an asset_id from your workspace, never both |
| Attachment type | input_image only; PDFs and other files are not available yet |
Which errors point at the image?
Three codes are specific to attachments. 400 invalid_attachment covers a wrong type, a missing or non-HTTPS URL, both image_url and asset_id on one item, or a source that is not an allowed image. 400 attachment_not_found means the asset_id is unknown in this workspace. 413 attachment_too_large means an image over 30 MB or a set over 500 MB. A fourth, 502 attachment_fetch_failed, means Sume could not fetch the URL: an unreachable host, hotlink protection, or a non-2xx answer. For that one, upload the image first and send its asset_id, or host it somewhere that serves it without checking a referrer.
How do images and output_schema work together?
They compose: the images reach the agent, and the run's output is parsed against your schema after the run completes. That is the shape for caption, alt text or tagging jobs where a database column wants a string and not a paragraph. The accepted call returns 202 with a run receipt, not a choices[] answer, and you poll status_url until the run reaches a terminal status. The create call is asynchronous because an agent turn may open a sandbox, call tools and generate media.
Two limits to plan around. Every completion runs in a fresh thread, and the docs list continuing a prior thread with thread_id as not available yet. And streaming and a synchronous OpenAI-compatible response are not available, so a plug-in that expects choices[] will not work unchanged.
URL or asset_id: which should you send?
Use image_url when the file is already public over HTTPS. The Format docs say Sume fetches it at create time, so it must be reachable without authentication, and it checks the real type and size and copies the image into durable storage. A broken or private image therefore fails the create with an error you can act on, instead of killing the run minutes later. The accepted types are JPEG, PNG, WebP, GIF and AVIF.
Use asset_id when the image sits behind your own auth or a hotlink rule. Upload it through the Assets API first, once it is ready in the same workspace, and send the id; an asset_id or a URL already on media.sume.com is not copied again. Send one or the other on an item, never both. An optional filename sets the label the agent sees, which defaults to the URL's basename; name files descriptively when you attach several, so the instruction can refer to packshot.png rather than "the second image".
Last, remember the Agent Completions spend cap. An image-only caption job rarely generates media, but the cap is still required on every request, so send generation_spend_cap_usd even for a read-only task.
Sources
Related posts
More in Developers
- AI image API 429s: queue_full vs rate_limited, and how to retry each
Sume returns 429 for two different reasons. rate_limited means back off; queue_full means wait for jobs to finish. A Python retry that treats them differently.
- Is there an asset library API for AI images and videos?
Sume has no folders or tags. Your library is completed jobs plus durable media.sume.com artifacts, which you list, label by Idempotency-Key, and download.
- AI image model fallback in Python: try the next model on a 502
Image models launch and fail on different days. A Python loop that tries the next Sume model id on 502 or 503, stops on 400, and keeps the 202 job path intact.
- Claude Cost Report API: daily buckets by workspace vs Sume /v1/usage
Anthropic's cost_report endpoint returns USD cost in 1d buckets, groupable by workspace or description. Sume's /v1/usage sums one thread, run or job instead.
Written by Sume