Which Sume spend guard applies: Completions, Formats, schedules, MCP

One table of the spend and retry guards on each Sume surface, for teams that put a new LLM or decision model in front of paid calls.

5 min readSume
All posts

Each Sume surface has its own guard, and they are not the same field. Agent Completions require generation_spend_cap_usd. Format runs report the same cap in usage. A schedule keeps a saved cap, and an API-triggered run is clamped to the smaller of the request and that cap. Hosted MCP paid tools take max_spend_usd, which Sume enforces only when you send it. A new LLM or decision model in front of any of these changes none of this.

Guards by surface

Use the table when you wire a new model into a flow, and check each row before the first paid call.

Spend and retry guards by surface (read 2026-10-05)
SurfaceSpend ceilingRetry safetyPreview
Agent Completionsgeneration_spend_cap_usd, requiredIdempotency-KeyNone, the cap is the limit
Format runsCap in the request, shown in usageIdempotency-KeyRead io first
Scheduled runsSaved cap, min(request, cap)Idempotency-Key on API triggerNone
Hosted MCP paid toolsmax_spend_usd, only when sentidempotency_keydry_run=true

What the receipts do not show

The receipt figure usage.billable_amount_usd_micros excludes the agent's own LLM turn, which bills the Agent wallet. GET /v1/usage is the billing record. Do not size an alert on the receipt alone.

Where a decision model fits

Put the model's answer upstream of all four, and keep every guard in the table on. A confident model is a reason to run, never a reason to remove a ceiling.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume