generation_admission_preview: check before paid hosted MCP calls
Sume's hosted MCP has generation_admission_preview, dry_run and max_spend_usd. What each one checks, when to use it, and when a plain single create is fine.

Before an expensive burst of paid MCP calls, call generation_admission_preview, or call the paid tool itself with dry_run=true. Both show whether the work would be admitted without submitting a job; max_spend_usd then caps what a real submit may spend. The docs say ordinary single creates do not need this extra step.
What are Sume's MCP safety gates?
Hosted MCP is read-only by default under OAuth mcp:read. Mutating and paid tools stay hidden until the session has mcp:write or uses an API key. The docs add that spend is governed by wallet and admission, and that there is no mcp:paid scope.
Three gates are about money and retries, listed on MCP tools and gates.
| Gate | Required? | What it does |
|---|---|---|
| idempotency_key | Required on write and paid tools | Stable key for transport and dedup; not human approval |
| dry_run=true | Optional | Admission and cost preview only; the job is not submitted |
| max_spend_usd | Optional | Enforced only when you provide it |
When should I preview and when can I skip it?
Preview before bursts: dozens of images, a batch of avatar videos, or any loop an agent runs unattended. The preview reads your balance, queue and concurrency state, so you learn about a 402 insufficient_credits or a full queue before you spend a turn on it.
Skip it for one ordinary create. The docs say plainly that single creates do not need "admission theater". The reservation at submit is the real gate, and an idempotent retry is safe.
What does a safe paid call look like?
The pattern on the docs' avatar playbook is: first call with dry_run=true and review the preview, then repeat with dry_run omitted or false to submit, then poll with jobs_status or jobs_wait and read jobs_result. Keep the same idempotency_key across the dry run and the submit only if the payload is identical; a new payload needs a new key.
The payload shape follows the tool's schema. Inspect it with tools_schema before you build it, as the docs advise.
{
"idempotency_key": "avatar-create-2026-10-02-001",
"dry_run": true,
"max_spend_usd": 2,
"payload": {
"avatar_handle": "studio_presenter",
"input": {
"type": "prompt",
"prompt": "A friendly studio presenter in neutral lighting"
}
}
}What do these gates not do?
max_spend_usd is enforced only when you pass it, so an agent that omits it has no per-call ceiling beyond the wallet. A preview is a snapshot: counts can change right after it, because other clients and workers move jobs at the same time. Neither replaces a spend cap on runs that call tools in a loop. For that, use the run-level cap described in the docs for Formats and Agent Completions.
Sources
Related posts
More in Pricing
- Can a per-run spend cap raise the limit? Formats yes, schedules no
Formats honor a per-run generation_spend_cap_usd above the Format cap, schedules clamp it, and Agent Completions require it. Defaults and edge cases.
- Rask AI minute credits and 3x lip-sync vs a Sume dubbing pipeline
Rask bills 1 credit per video minute for Standard lip-sync, 3 for Enhanced. Sume has no dubbing endpoint; a chained pipeline runs about $0.55-$1.55 per 10 min.
- Realtime voice API cost per hour: Grok, GPT-Live-1, Gemini Live
Per-hour cost of realtime voice APIs from the vendors' own pages, set against Sume's async STT plus TTS jobs, with the arithmetic and its assumptions shown.
- Recraft API pricing: 1,000 API units per dollar, V4.1 Flash $0.007
Recraft sells prepaid API units at $1.00 for 1,000. V4.1 Flash is 7 units ($0.007) per raster image, V4.1 is 35 units ($0.035), V4.1 Pro is 210 units ($0.21).
Written by Sume