Alt text for a 30-image gallery in one Sume Agent Completion
One Agent Completion call can take up to 30 images and return an alts array under an object schema. Python stdlib script with the cap, poll and limits.

One POST /v1/agent/completions call can look at up to 30 images and return one alt text per image, as long as the schema's root is an object. A top-level array is rejected, so the list goes under a key such as alts. The call is asynchronous: it answers 202 with a receipt, you poll, and a completed run holds output in the shape you asked for, or output: null with a reason. It never hands back a guess that fails your schema.
This is a different tool from a Format bulk queue. A bulk queue runs one saved Format many times. A completion is one ad-hoc agent task, so a whole gallery for one product is a single run with a single spend cap, rather than 30 runs.
The script
Pass image URLs as arguments. Set SUME_API_KEY, which needs the agent_completions:read and agent_completions:write scopes. A key made before Agent Completions shipped lacks them and gets 403 insufficient_scope; a service-account key cannot create completions at all.
import asyncio, json, os, sys, urllib.request
BASE = os.environ.get("SUME_BASE", "https://api.sume.com/v1")
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}
def call(url, body=None):
data = json.dumps(body).encode() if body else None
req = urllib.request.Request(url, data=data, headers=HEAD)
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)["data"]
async def main(urls):
schema = {"type": "object", "additionalProperties": False, "required": ["alts"],
"properties": {"alts": {"type": "array", "items": {"type": "string"}}}}
parts = [{"type": "input_text", "text": f"Write one alt text under 125 characters for each of "
f"these {len(urls)} product images, in the order sent."}]
parts += [{"type": "input_image", "image_url": u} for u in urls]
run = call(BASE + "/agent/completions", {
"messages": [{"role": "user", "content": parts}],
"output_schema": {"name": "alts", "schema": schema}, "generation_spend_cap_usd": 2})
while call(run["status_url"])["next_action"] == "poll_status":
await asyncio.sleep(10)
done = call(f"{BASE}/agent-runs/{run['id']}")
print(done["status"], json.dumps(done["output"], indent=1))
asyncio.run(main(sys.argv[1:]))The rules it follows
generation_spend_cap_usd has no default on this endpoint. Leave it out and the request fails with 400 invalid_request, because an unattended agent has no approval prompt and the cap stands in for it. Pick the highest amount you accept for the one run. The script uses 2.
The images are input_image content parts, with an input_text part first that tells the agent what to do. The same item shape is accepted at the top level in attachments; Sume merges both sources into one list. The limit is 30 images, 30 MB each and 500 MB in total. A bad item gives 400 invalid_attachment, an unknown asset_id gives 400 attachment_not_found, a size breach gives 413 attachment_too_large, and a host that cannot be fetched gives 502 attachment_fetch_failed.
Polling follows next_action: the loop keeps going while it is poll_status, then reads the full run from GET /v1/agent-runs/{id}. The schema has additionalProperties: false on the root and lists every property in required, as the structured-output rules demand for every object.
| Item | Rule | If broken |
|---|---|---|
| Images per call | At most 30 | 400 invalid_attachment or invalid_request |
| Size | 30 MB per image, 500 MB total | 413 attachment_too_large |
| Spend cap | Required, no default | 400 invalid_request |
| Schema root | Must be an object | 400 output_schema_invalid, violation root_must_be_object |
| Message roles | system and user only | assistant turns are rejected |
Matching alt text to images
The prompt asks for the alt texts in the order sent, and the schema returns a plain array, so the first string belongs to the first URL. The schema cannot enforce an array length, since only minItems and maxItems are supported, and it does not tie each string to a URL. Check len(alts) == len(urls) yourself before you write anything to a catalog, and treat a mismatch as a failed run. If you need a stronger tie, make each array item an object with the URL echoed back and compare it.
Each completion starts a fresh thread. You cannot continue one with thread_id, so a second call about the same gallery re-attaches the images.
Sources
Related posts
More in Developers
- API key scopes for Sume: which key can call which endpoint family?
Sume API keys carry fixed scopes: formats:write, actions:read, agent_completions:write, account:read. Which scope each route needs, and why old keys get a 403.
- Arabic speech to text API: Sume STT with language_code ar
Transcribe Arabic audio with Sume STT: send language_code ar, check the reported language, and review the text. $0.01 per audio minute.
- Avatar job tracking table: which Sume ids to store and why
Avatar work produces a handle, a job id, a preview id and a video id. A small SQL table that keeps them straight, plus the status fields to poll.
- Avatar video webhook mode: what arrives and what to poll anyway
Use mode webhook for a Sume avatar video and Sume posts one terminal event: completed, failed or canceled. Payload, signature headers and the polling backup.
Written by Sume