Agent Completions 400: generation_spend_cap_usd is required
POST /v1/agent/completions returns 400 invalid_request without generation_spend_cap_usd. It has no default: why, how to size it, and what the receipt echoes.

Sume's POST /v1/agent/completions returns 400 invalid_request when the body has no generation_spend_cap_usd. The field has no default and no environment fallback, so every call must say the most it may spend on that one run. Add it as a number of US dollars and resend.
The same code is also returned for other reasons, so read the error details before assuming it is the cap. Both cases are covered below, along with how to size the number.
Why is the spend cap required and not defaulted?
The Agent Completions docs are explicit about the reason. A completion is an unattended agent with tools and access to your generation wallet. In the chat UI a person approves spend interactively; a backend caller cannot do that. The cap is the substitute, so it is mandatory and Sume does not pick a number for you.
Set it per run to the most you are willing to lose on that single task. A cap that is a constant in your config and never changes is a cap sized for the biggest job, which is the wrong ceiling for the small ones.
What does a correct request look like?
Send exactly one of instruction or messages, plus the cap. An Idempotency-Key header is optional but worth sending: replaying a key returns the original receipt with idempotency_hit: true, and reusing a key with a different payload returns 409 idempotency_conflict.
curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: launch-hero-2026-10-02" \
-d '{
"instruction": "Make one hero image for the autumn sale page.",
"generation_spend_cap_usd": 2
}'What else returns 400 invalid_request?
The cap is one of several causes. The docs list these under the same code:
| Cause | Fix |
|---|---|
| Missing generation_spend_cap_usd | Add the cap |
| Neither or both of instruction and messages | Send exactly one |
| An assistant turn in messages | Remove it; every completion runs in a fresh thread |
| Malformed input | Send a JSON object |
| A model other than sume-agent | Omit model, or send sume-agent |
How do you size the number?
The docs point to the metered rates on the API pricing page, because the cap is a ceiling against those rates. Estimate what one run will produce (images, seconds of video, audio) from that page, add headroom for a retry inside the run, and round up. Do not copy the 2 or 5 in the docs' examples; they are illustrations.
The accepted receipt echoes the number back as usage.generation_spend_cap_usd_micros, in millionths of a dollar: the docs' example shows 5000000 for a five-dollar cap. Log it with the run id so a later bill review can compare cap and actual spend, which is recorded in usage on the finished run.
What the docs do not say
The Agent Completions page does not describe what the run does when it reaches the cap, so do not design around a particular behaviour such as a clean partial result. Treat the cap as the limit you set, and test it with a small value on a throwaway task before you rely on it. Keys created before the feature carry no agent_completions:* scopes and fail with 403 insufficient_scope; service-account keys cannot create completions at all.
How does the cap fit with retries and webhooks?
Send an Idempotency-Key on every create from a queue worker. If the network drops after Sume accepts the request, a retry with the same key returns the original receipt with idempotency_hit: true instead of starting a second run with a second cap. A different payload under the same key returns 409 idempotency_conflict, which usually means the key was derived from something too coarse, such as the date alone.
To learn the outcome without polling, set communication.webhook_url to a public HTTPS URL; Sume notifies it when the run reaches a terminal status (Run webhooks). Otherwise poll status_url from the receipt until next_action stops being poll_status, and read the recorded spend from usage on the finished run.
A short checklist before you ship the integration: a new API key created after Agent Completions shipped, with agent_completions:write and agent_completions:read; a per-task cap computed from the pricing page; an idempotency key per business intent; and a log line pairing run id, cap and recorded spend.
Sources
Related posts
More in Developers
- Agent Completions input_image: send a photo in messages[]
Send an image to Sume's Agent Completions as an input_image content part inside messages[], merge it with attachments, and get a caption back as typed output.
- AI image API 429s: queue_full vs rate_limited, and how to retry each
Sume returns 429 for two different reasons. rate_limited means back off; queue_full means wait for jobs to finish. A Python retry that treats them differently.
- Is there an asset library API for AI images and videos?
Sume has no folders or tags. Your library is completed jobs plus durable media.sume.com artifacts, which you list, label by Idempotency-Key, and download.
- AI image model fallback in Python: try the next model on a 502
Image models launch and fail on different days. A Python loop that tries the next Sume model id on 502 or 503, stops on 400, and keeps the 202 job path intact.
Written by Sume