Stop an AI video agent overspending: the four Sume spend gates
An unattended Sume agent can only spend what four gates allow: the wallet, a per-run cap, a schedule ceiling, and a Format cap. Values and failure codes inside.

Four gates limit what an unattended Sume video agent can spend: the workspace wallet (checked when the run is created), the per-run generation_spend_cap_usd, the ceiling stored on a schedule, and the cap stored on a Format. A run ends as failed when it tries to go past its effective cap, so the worst case of any one run is a number you chose before it started.
The gates differ by surface, so the table below lists each one with the value and the failure you will see. Read it before you wire an agent to a cron or a queue, because an unattended caller never sees the spend-approval prompt that the chat UI shows a person.
The gates side by side
Every number below comes from the Sume docs pages in the sources list.
| Surface | Cap field | Default | What happens past the limit |
|---|---|---|---|
| Agent Completion | generation_spend_cap_usd | None. Required | Missing cap is 400 invalid_request |
| Scheduled run | Cap stored on the schedule; per-run override | $1.00 if unset | Override clamps to the lower of request and schedule cap; 0 is rejected; null drops the automation ceiling |
| Format run | Format cap, or generation_spend_cap_usd on the request | $400 if the Format never set one | 0 or above 500 is 400; null means the platform maximum of $500 |
| Any run, at create | Workspace wallet | Balance | 402 insufficient_credits with next_action: add_funds |
Gate 1: the wallet
The wallet is checked when you create the run. If it cannot fund the run you get 402 insufficient_credits (or 402 organization_wallet_not_provisioned for an organization workspace nobody has funded) and nothing ran. Retrying before you add funds returns the same answer, so a scheduler should alert on this code instead of looping.
Gate 2: the cap you send
For Agent Completions the cap is mandatory on purpose. The chat UI protects you with an interactive approval prompt, a backend caller does not get one, and the docs say the cap replaces the prompt. Set it per run to the most you accept for that task.
For Format runs the request value is accepted even when it is higher than the Format's own cap, up to 500, so a careless client can raise the ceiling. Keep the per-run number in your own config and review it like any other spending limit.
Gate 3: the schedule ceiling
A schedule stores a per-run cap. A caller of the API trigger can lower it with a number, and it can never raise it above the schedule. Sending null removes the automation ceiling, although wallet balance, generation admission and organization limits still apply, so treat null as a review-required value.
Gate 4: the Format cap, and what the receipt tells you
Every receipt carries usage.generation_spend_cap_usd_micros and usage.billable_amount_usd_micros. When a run fails with the generic format_run_failed, compare those two numbers first, because the docs state that a run which tried to spend more than its cap gets that code.
Two limits on reading usage. It covers generation spend, not the agent's own model turn, which bills the separate Agent wallet. It is also a receipt figure, not an invoice, so reconcile against GET /v1/usage.
A batch multiplies the gates
A bulk queue has no queue-level limit, so the exposure is items times the per-item cap. One hundred items at a $120 cap is $12,000 of worst-case exposure for one queue. Send a smaller cap per item, and keep the batch size inside what you would approve by hand.
Sources
Related posts
More in Agents
- MCP wants a human in the loop: Sume gates for unattended agents
MCP says a human SHOULD be able to deny tool calls. For unattended agents, Sume adds scopes, idempotency keys and a required Agent Completions cap.
- Auditing agent tool calls: MCP log advice, Sume script_run journal
The MCP spec tells clients to log tool usage for audit. For Sume's script_run, the response includes a calls[] journal and child jobs[] to use with jobs_wait.
- Mistral Large 4 preview: Sume's agent model field takes one value
Mistral Large 4 is reported in preview. Sume lists no Mistral Large id; Agent Completions accepts only sume-agent. What to send instead, and what it bills.
- One webhook handler for Sume agent, Format and schedule runs
A single receiver can handle all three Sume run types: verify the HMAC over timestamp.body, dedupe on request_id, and branch on outcome (ok, degraded, error).
Written by Sume