A decision model's confidence is not a spend cap: Sume's real gates
Strands Decider and Clef return confidence scores. Only the Sume gates in this table, from idempotency_key to the required spend cap, limit what a run can cost.

No. A confidence score says how sure a model is about an answer. It does not limit what the next call can cost. The Strands Decider model card describes calibrated confidence scores, and it also lists limits: calibration is one temperature per primitive, and the model inherits label noise from public training data. Treat the number as advice and let the Sume gates do the limiting.
Which gate stops what
Each Sume gate stops a different failure. The decision model is the only row in the table that is advisory.
| Layer | Stops | Required? | Where it is documented |
|---|---|---|---|
| Decision model confidence | Items that probably do not need a run | No, advisory | Strands and Cloudflare pages |
idempotency_key | A retry that pays twice | Yes on MCP write and paid tools | MCP tools and gates |
dry_run=true | Surprise at the estimate | No | MCP tools and gates |
max_spend_usd | A call above your ceiling | No, enforced only when sent | MCP tools and gates |
generation_spend_cap_usd | A run above its ceiling | Yes on Agent Completions | Agent Completions |
| Schedule cap | A run above the saved cap | Default $1.00 if unset | Scheduled |
Why the cap is yours to set
On hosted MCP, max_spend_usd is optional, and Sume enforces it only when you provide it. A flow that trusts a high confidence score and omits it has no ceiling left except the wallet and admission. Send it on every paid call, and size it from the dry run, not from the model's certainty.
On Agent Completions the cap is mandatory because an unattended agent has tools and access to your generation wallet. On a schedule, the API clamps a per-run number to min(request, schedule cap): a request can lower the cap and can never raise it.
Use the score for triage, not for limits
Log the decision next to the receipt so that you can compare the two later. A decision with a high score on a run that hit its cap is worth a look. A decision with a low score on a run that finished cheaply is worth a look as well.
Sources
Related posts
More in Agents
- DeepSeek V4.1 Flash sees images: hand one to a Sume Agent Completion
V4.1 Flash is described as natively multimodal. To act on an image with Sume, pass it as an input_image attachment on an Agent Completion, up to 30 per run.
- Does the Sume run spend cap include the LLM turn? No, only generation
On a Sume run receipt, billable_amount_usd_micros is generation spend only; the agent's LLM turn bills a separate Agent wallet. How to budget both.
- Fire a Sume schedule from a decision, not a clock: API trigger
Set api_trigger_enabled on a Sume schedule, let Clef or Decider decide when it is worth running, then POST /v1/actions/{action_id}/runs with an idempotency key.
- Five generate_image calls in one turn: Sume MCP create budget
Hosted Sume MCP gives paid creates their own budget (20 per principal, 64 per process) and queues a call up to 20 seconds before wait_busy.
Written by Sume