Send product photos to the Sume agent API, get edited images back
Agent Completions takes up to 30 image attachments and a required spend cap, then returns generated images in output.images. When to use it over /v1/images.

To have an agent look at your product photos and make new images, call POST /v1/agent/completions with the photos as attachments (or input_image parts), a plain instruction, and a generation_spend_cap_usd. The call returns 202 and a receipt; the finished run puts the generated pictures in output.images.
Use it when the task changes on every call and you want the agent to choose the model and steps. Use POST /v1/images instead when you already know the model and want one predictable image per request.
Which route fits
The Agent Completions page separates three surfaces by how often the task changes. The Image API is a fourth, narrower path: one model, one request, one response.
| Route | Use it when | Response |
|---|---|---|
| POST /v1/images | You know the model and the inputs | 200 with image URLs, or 202 job |
| POST /v1/agent/completions | The task changes each call and the agent should decide | 202 receipt, then poll the run |
| Format run | A saved workflow, only inputs change | Run receipt |
The request
Attachments take up to 30 images. Each must be a public HTTPS URL or an asset_id, not both; an image over 30 MB or a set over 500 MB is refused with 413 attachment_too_large, and a host that cannot be reached returns 502 attachment_fetch_failed. The spend cap has no default: leaving it out returns 400 invalid_request. The docs explain why: a backend caller gets no approval prompt, so the cap takes its place.
The sample sends two photos and a short brief. Set the cap to the most you will accept for this one run.
import os
import httpx
body = {
"messages": [{"role": "user", "content": [
{"type": "input_text", "text": "Make a 4:5 lifestyle image of this product on a kitchen counter."},
{"type": "input_image", "image_url": "https://example.com/front.jpg"},
{"type": "input_image", "image_url": "https://example.com/side.jpg"},
]}],
"generation_spend_cap_usd": 1.5,
}
headers = {
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "listing-4412-lifestyle-1",
}
r = httpx.post("https://api.sume.com/v1/agent/completions", json=body, headers=headers, timeout=60)
r.raise_for_status()
print(r.status_code, r.json()["data"]["status_url"])Reading the result
Poll the status_url until next_action is no longer poll_status. A completed run fills output: the agent's last text is in output.text, and generated media is in output.images, output.videos, output.audio and output.files. Media URLs are durable media.sume.com HTTPS URLs. usage records the spend, so you can compare it with the cap you set.
A repeated call with the same Idempotency-Key returns the original receipt with idempotency_hit: true, and the same key with a different payload returns 409 idempotency_conflict. Use a key per job in your own system, such as a listing id plus a round number, so a retry after a network error does not start a second run.
Writing the brief
The agent chooses its own tools, so the brief is where you control the result. State the deliverable in one sentence, the count, the ratio and what must not change. For example: three images, 4:5, keep the label text as in the photo, no people. Attach only the photos that matter, since every attachment is something the agent has to look at.
The receipt echoes your cap in usage.generation_spend_cap_usd_micros, in millionths of a dollar, so a cap of 1.5 appears as 1500000. Size the cap from the rates on the API pricing page, and leave room for retries inside the run, because the agent may generate more than once before it settles on the images it returns.
Gotchas before you ship it
Two permissions catch teams out. Keys created before Agent Completions shipped lack the agent_completions:read and agent_completions:write scopes, so they fail with 403 insufficient_scope and you cannot add scopes to an existing key. Service-account keys cannot create completions at all.
- Create a new API key for this route and rotate to it.
- Send
assistantturns and the call fails: Sume rejects them, because each completion starts a new thread. - Streaming and a synchronous chat-style reply are not available; the call is asynchronous by design.
- Set
communication.webhook_urlto a public HTTPS URL if you would rather be called than poll.
Sources
Related posts
More in Agents
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
- Safe automation for AI agents that call paid APIs
Keep agents read-only by default, keep secrets out of logs, and on hosted MCP send an idempotency_key, preview with dry_run, and cap with max_spend_usd.
- Scheduled AI video agent runs: cron, API triggers, and receipts
A Sume schedule is a saved Agents automation that runs on a cron cadence and returns a run receipt. Author it in the dashboard; start and monitor runs by API.
- What is a video agent? How Sume defines and runs one
In Sume's docs, a video agent is a sandbox Agent that composes generation tools into a post-ready video. Brief it in chat, or call it over HTTP.
Written by Sume