generation_admission_preview: check before paid hosted MCP calls

Sume's hosted MCP has generation_admission_preview, dry_run and max_spend_usd. What each one checks, when to use it, and when a plain single create is fine.

4 min readSume
All posts

Before an expensive burst of paid MCP calls, call generation_admission_preview, or call the paid tool itself with dry_run=true. Both show whether the work would be admitted without submitting a job; max_spend_usd then caps what a real submit may spend. The docs say ordinary single creates do not need this extra step.

What are Sume's MCP safety gates?

Hosted MCP is read-only by default under OAuth mcp:read. Mutating and paid tools stay hidden until the session has mcp:write or uses an API key. The docs add that spend is governed by wallet and admission, and that there is no mcp:paid scope.

Three gates are about money and retries, listed on MCP tools and gates.

Paid-call gates on hosted MCP, from the Sume docs, read 2026-10-02
GateRequired?What it does
idempotency_keyRequired on write and paid toolsStable key for transport and dedup; not human approval
dry_run=trueOptionalAdmission and cost preview only; the job is not submitted
max_spend_usdOptionalEnforced only when you provide it

When should I preview and when can I skip it?

Preview before bursts: dozens of images, a batch of avatar videos, or any loop an agent runs unattended. The preview reads your balance, queue and concurrency state, so you learn about a 402 insufficient_credits or a full queue before you spend a turn on it.

Skip it for one ordinary create. The docs say plainly that single creates do not need "admission theater". The reservation at submit is the real gate, and an idempotent retry is safe.

What does a safe paid call look like?

The pattern on the docs' avatar playbook is: first call with dry_run=true and review the preview, then repeat with dry_run omitted or false to submit, then poll with jobs_status or jobs_wait and read jobs_result. Keep the same idempotency_key across the dry run and the submit only if the payload is identical; a new payload needs a new key.

The payload shape follows the tool's schema. Inspect it with tools_schema before you build it, as the docs advise.

{
  "idempotency_key": "avatar-create-2026-10-02-001",
  "dry_run": true,
  "max_spend_usd": 2,
  "payload": {
    "avatar_handle": "studio_presenter",
    "input": {
      "type": "prompt",
      "prompt": "A friendly studio presenter in neutral lighting"
    }
  }
}

What do these gates not do?

max_spend_usd is enforced only when you pass it, so an agent that omits it has no per-call ceiling beyond the wallet. A preview is a snapshot: counts can change right after it, because other clients and workers move jobs at the same time. Neither replaces a spend cap on runs that call tools in a loop. For that, use the run-level cap described in the docs for Formats and Agent Completions.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume