Which Sume spend guard applies: Completions, Formats, schedules, MCP
One table of the spend and retry guards on each Sume surface, for teams that put a new LLM or decision model in front of paid calls.

Each Sume surface has its own guard, and they are not the same field. Agent Completions require generation_spend_cap_usd. Format runs report the same cap in usage. A schedule keeps a saved cap, and an API-triggered run is clamped to the smaller of the request and that cap. Hosted MCP paid tools take max_spend_usd, which Sume enforces only when you send it. A new LLM or decision model in front of any of these changes none of this.
Guards by surface
Use the table when you wire a new model into a flow, and check each row before the first paid call.
| Surface | Spend ceiling | Retry safety | Preview |
|---|---|---|---|
| Agent Completions | generation_spend_cap_usd, required | Idempotency-Key | None, the cap is the limit |
| Format runs | Cap in the request, shown in usage | Idempotency-Key | Read io first |
| Scheduled runs | Saved cap, min(request, cap) | Idempotency-Key on API trigger | None |
| Hosted MCP paid tools | max_spend_usd, only when sent | idempotency_key | dry_run=true |
What the receipts do not show
The receipt figure usage.billable_amount_usd_micros excludes the agent's own LLM turn, which bills the Agent wallet. GET /v1/usage is the billing record. Do not size an alert on the receipt alone.
Where a decision model fits
Put the model's answer upstream of all four, and keep every guard in the table on. A confident model is a reason to run, never a reason to remove a ceiling.
Sources
Related posts
More in Developers
- Which Sume video models take 30 seconds? Filter /v1/videos/models
Seedance 2.5 and Wan 3.0 list 30 s in supported_durations. A 12-line Python script reads GET /v1/videos/models and prints every id that accepts your length.
- Which Sume video rows take video_url: Omni edit, Recast or Genjutsu
Three Sume rows take a source video_url: Gemini Omni edit ($0.125 a second at 720p), H3 Max Recast ($0.375 and up) and Genjutsu ($0.3975 and up). Pick by task.
- Which URL goes into video-captions after a trim: the new artf_
After video-trim, caption the new video_url from the trim result, not the source render. It is a new artifact on media.sume.com and a public HTTPS URL.
- Which video model takes 30 seconds? Filter the Sume catalog in Python
Ask GET /v1/videos/models which models accept duration 30 at 1080p instead of hard-coding ids. A short Python filter, and why Omni and Kling 3 drop out.
Written by Sume