Get the edited image URL as a named field from an Agent Completion
Attach the photo and an output_schema in one POST /v1/agent/completions call. Sume parses output into your own fields, such as edited_url and changes.

Send the photo in attachments and an output_schema in the same POST /v1/agent/completions call. The images go to the agent, and when the run completes Sume parses output against your schema, so you can ask for fields like edited_url and changes instead of scraping text.
What the docs promise
The Agent Completions page says attachments and output_schema can be used together: the images go to the agent, and after the run completes Sume still parses the output against your schema (Agent Completions). A turn that only has images is allowed. Without a text part, the agent is told to use the attached files.
| Field | Required | Role |
|---|---|---|
| instruction or messages | one of them | what to change |
| attachments | no | up to 30 images, 30 MB each |
| output_schema | no | your own JSON shape for output |
| generation_spend_cap_usd | yes | most you accept for the run |
| Idempotency-Key header | no | same key returns the original receipt |
A schema for an edit
Ask for the URL and a short list of what changed. The agent fills the URL from the media it generated, and your code checks it before it uses it.
payload = {
"instruction": "Replace the grey wall with warm white. Keep everything else.",
"attachments": [{"type": "input_image", "image_url": "https://example.com/room.jpg"}],
"output_schema": {
"name": "edit_result",
"schema": {
"type": "object",
"properties": {
"edited_url": {"type": "string"},
"changes": {"type": "array", "items": {"type": "string"}},
},
"required": ["edited_url", "changes"],
"additionalProperties": False,
},
},
"generation_spend_cap_usd": 1,
}Check what you get back
Poll GET /v1/agent-runs/{id} until the run is completed. The generated files are also in output.images, so compare the URL in your field with that list. A URL in edited_url that is not in output.images is a sign the agent wrote something it did not generate, and you should treat the run as failed.
When not to use it
If all you need is one URL back, the Image API already returns data[].url and needs no schema. A schema is useful when the task has more than one output, for example an edited image, a list of changes and a flag for whether the logo stayed untouched. Remember that the agent's own report of a change is not proof of it: diff the images before you trust the changes list.
Handling the failure cases
A run can finish as failed or canceled, and then your schema does not apply. Check the status first. If a run fails after spending on a generation, the cap is what bounds the loss, so choose it for the task and not as a round number.
Send an Idempotency-Key so a client retry after a timeout gets the original receipt and does not start a second run. If you send the same key with a different payload, you get a 409 and should fix the caller.
Sources
Related posts
More in Developers
- Agent Completion output_schema shaped like a Clef or Decider answer
Map Clef and Strands Decider answer types (yes/no, choice, score) onto a Sume output_schema with enum and integer bounds, and see what it does not check.
- Agent quoted $13.31 for a 13-cent Sume job: the micros divisor
Divide billable_amount_usd_micros by 1,000,000 for dollars. Dividing by 10,000 gives cents and overstates spend 100x. Sume MCP flags this as usd_unit_mismatch.
- catalog_list vs tools_list: finding HTTP-only Sume features
tools_list shows what this MCP session can call. catalog_list shows API capabilities, some with no MCP tool. Read both before telling a user it can't be done.
- Agent run webhook: created_at orders deliveries, request_id dedupes
An agent.run.terminal delivery has two ids that look alike. request_id dedupes retries; created_at orders deliveries. Includes a Python receiver.
Written by Sume