max_spend_usd on Sume MCP: floors to micros, missing estimate fails

max_spend_usd takes 0 to 10,000, floors to micros, and runs on the preview before dry_run returns. No usable estimate means missing_usage_estimate, not a pass.

4 min readSume
All posts

On Sume's hosted MCP, max_spend_usd accepts 0 to 10,000, is floored to integer micros, and is checked against the admission preview before the dry-run result returns and before any submit. If the preview has no integer billable_amount_usd_micros, the call fails closed with missing_usage_estimate; it does not slip through uncapped.

What the check does

Order matters. The server prices the request with the generation admission preview, then compares that estimate to your cap, and only then decides whether to return a dry-run object or submit. So the cap protects the dry run path and the real path identically.

Your dollar cap is converted to micros by flooring. A cap of 0.1234569 dollars becomes 123,456 micros, so an estimate of 123,457 micros is refused even though it rounds to the same cents. Send a cap you are happy to be strict about.

max_spend_usd outcomes on paid Sume MCP tools (read 2026-10-05 against the Sume codebase)
SituationError codeFields returned
Estimate above the capmax_spend_exceededusage fields, max_spend_usd, preview
Preview has no integer micros estimatemissing_usage_estimatefails closed
Preview would not accept the requestgeneration_admission_rejectedpreview
Cap outside 0 to 10,000input validation errorrange

Why fail closed matters for agents

A guard that silently passes when it cannot measure is worse than no guard, because the agent believes it is capped. Sume treats an unmeasurable estimate as a refusal. Your retry policy should do the same: on missing_usage_estimate, do not remove the cap and try again. Read the preview, fix the request, or ask a human.

A model's confidence is a different thing entirely. The decision model confidence post explains why a yes/no gate from a small model should sit in front of the cap, never replace it.

Client-side mirror

Mirror the floor in your own pre-check so your logs match the server's.

import math

def cap_to_micros(max_spend_usd: float) -> int:
    if not 0 <= max_spend_usd <= 10_000:
        raise ValueError('max_spend_usd must be 0..10000')
    return math.floor(max_spend_usd * 1_000_000)

def allowed(estimate_micros, cap_usd: float) -> bool:
    if not isinstance(estimate_micros, int):
        return False  # fail closed, like missing_usage_estimate
    return estimate_micros <= cap_to_micros(cap_usd)

print(allowed(123_457, 0.1234569), allowed(None, 5))

Limits

The cap is optional on most paid tools, so an agent that never sends it is not capped by this mechanism; send it on every paid call. Two dev-only analysis tools require it. Fixed-price tools use a separate fixed-price check. The cap compares against the estimate, not the final debit, so read the usage ledger after the job completes.

Checklist before you ship

  • Send max_spend_usd on every paid call, not only on dry runs.
  • Treat missing_usage_estimate as a stop, never as a reason to drop the cap.
  • Floor your own cap to micros before comparing it with a logged estimate.
  • After completion, compare the cap with the debited amount in the usage ledger.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume