Stop an AI video agent overspending: the four Sume spend gates

An unattended Sume agent can only spend what four gates allow: the wallet, a per-run cap, a schedule ceiling, and a Format cap. Values and failure codes inside.

5 min readSume
All posts

Four gates limit what an unattended Sume video agent can spend: the workspace wallet (checked when the run is created), the per-run generation_spend_cap_usd, the ceiling stored on a schedule, and the cap stored on a Format. A run ends as failed when it tries to go past its effective cap, so the worst case of any one run is a number you chose before it started.

The gates differ by surface, so the table below lists each one with the value and the failure you will see. Read it before you wire an agent to a cron or a queue, because an unattended caller never sees the spend-approval prompt that the chat UI shows a person.

The gates side by side

Every number below comes from the Sume docs pages in the sources list.

Spend controls by surface (read 2026-10-07)
SurfaceCap fieldDefaultWhat happens past the limit
Agent Completiongeneration_spend_cap_usdNone. RequiredMissing cap is 400 invalid_request
Scheduled runCap stored on the schedule; per-run override$1.00 if unsetOverride clamps to the lower of request and schedule cap; 0 is rejected; null drops the automation ceiling
Format runFormat cap, or generation_spend_cap_usd on the request$400 if the Format never set one0 or above 500 is 400; null means the platform maximum of $500
Any run, at createWorkspace walletBalance402 insufficient_credits with next_action: add_funds

Gate 1: the wallet

The wallet is checked when you create the run. If it cannot fund the run you get 402 insufficient_credits (or 402 organization_wallet_not_provisioned for an organization workspace nobody has funded) and nothing ran. Retrying before you add funds returns the same answer, so a scheduler should alert on this code instead of looping.

Gate 2: the cap you send

For Agent Completions the cap is mandatory on purpose. The chat UI protects you with an interactive approval prompt, a backend caller does not get one, and the docs say the cap replaces the prompt. Set it per run to the most you accept for that task.

For Format runs the request value is accepted even when it is higher than the Format's own cap, up to 500, so a careless client can raise the ceiling. Keep the per-run number in your own config and review it like any other spending limit.

Gate 3: the schedule ceiling

A schedule stores a per-run cap. A caller of the API trigger can lower it with a number, and it can never raise it above the schedule. Sending null removes the automation ceiling, although wallet balance, generation admission and organization limits still apply, so treat null as a review-required value.

Gate 4: the Format cap, and what the receipt tells you

Every receipt carries usage.generation_spend_cap_usd_micros and usage.billable_amount_usd_micros. When a run fails with the generic format_run_failed, compare those two numbers first, because the docs state that a run which tried to spend more than its cap gets that code.

Two limits on reading usage. It covers generation spend, not the agent's own model turn, which bills the separate Agent wallet. It is also a receipt figure, not an invoice, so reconcile against GET /v1/usage.

A batch multiplies the gates

A bulk queue has no queue-level limit, so the exposure is items times the per-item cap. One hundred items at a $120 cap is $12,000 of worst-case exposure for one queue. Send a smaller cap per item, and keep the batch size inside what you would approve by hand.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume