Agent Completions request limits: 100,000 characters, 50 messages
The Sume Agent Completions schema caps instruction at 100,000 characters, messages at 50, and images at 30. Know where each limit is enforced before you send.

Most Agent Completions problems come from a missing spend cap, but the request schema holds a few size limits that are worth knowing before you build a pipeline. They live in the public request schema for POST /v1/agent/completions, and breaking one gives a 400 invalid_request, not a failed run.
The limits in the schema
instruction is a string of at most 100,000 characters. messages needs between 1 and 50 items, and each item needs a role of system or user, plus content. The content can be a string, or an array of parts with a type of text, input_text or input_image. An image part takes image_url, a public HTTPS URL of at most 2,048 characters, or an asset_id, and a filename of at most 180 characters.
model, when present, must be sume-agent. primary_output_key is at most 64 characters. input must be a JSON object. The call takes either instruction or messages, not both, and generation_spend_cap_usd is required and cannot be 0.
| Field | Limit |
|---|---|
| instruction | 100,000 characters |
| messages | 1 to 50 items, roles system or user |
| image_url | Public HTTPS, up to 2,048 characters |
| images per request | 30, from attachments and input_image parts together |
| image size | 30 MB each, 500 MB total (413 attachment_too_large) |
| primary_output_key | 64 characters |
Images are fetched at create time
The API fetches each image and copies it into Sume storage while the call is open. An unreachable or oversized image therefore fails the create call, not the run later. That makes the failure cheap and early. It also means a slow image host can slow the 202. Host images on a fast, public HTTPS origin, or upload them first and send an asset_id.
Split work, not the schema
If you hit the message cap, you are probably replaying a long chat. Messages are flattened into one prompt in this version, so the number of turns matters less than their text. Summarize older turns into the instruction, or move bulk data to input, where it is written to a file the agent reads.
The limits describe one request. A run is not a conversation store: send what the run needs, and keep your own history.
Sources
Related posts
More in Agents
- Preflight a paid Sume MCP call: balance_get, then dry_run
Before a paid Sume tool call, an agent should read the balance and run the tool with dry_run=true. Then submit with a fresh idempotency_key and an optional cap.
- API-triggered scheduled run: key per week, not $(uuidgen)
The docs' example sends Idempotency-Key: $(uuidgen), which changes on every retry. Derive the key from the week so a retry returns the same run.
- ChatGPT confirms write tools; Sume MCP also gates them by scope
ChatGPT developer mode asks before write actions. Sume's hosted MCP adds a second gate: OAuth mcp:write, an idempotency_key and an optional max_spend_usd.
- Claude 'Allow always' on a paid Sume tool: what still caps spend
Claude custom connectors let you approve a tool once and keep approving it. For a paid Sume tool, the scope, idempotency key and max_spend_usd still apply.
Written by Sume