Preflight a paid Sume MCP call: balance_get, then dry_run

Before a paid Sume tool call, an agent should read the balance and run the tool with dry_run=true. Then submit with a fresh idempotency_key and an optional cap.

5 min readSume
All posts

Run three steps in order: read the balance, preview the call, then submit. Sume's hosted MCP has balance_get as a free read, dry_run=true on paid tools as a cost and admission preview that does not submit a job, and an optional max_spend_usd that Sume enforces only when you send it (read 2026-10-05 in the repo docs). A paid submit also needs an idempotency_key.

Why a normal single call can skip it

The docs are careful here. They recommend generation_admission_preview or dry_run before expensive bursts, and say a normal single create does not need these admission steps. The preview earns its cost when an agent is about to fan out, or when a person has not confirmed spend.

There is also no mcp:paid OAuth scope. Under mcp:read the paid tools are hidden or return insufficient_scope; under mcp:write or an API key they are available and spend is decided by the wallet and admission. So the gates in the request are the control you have.

Preflight steps and what each costs, read 2026-10-05
StepTool or fieldSpends money?
Check walletbalance_getNo
Preview a burstgeneration_admission_previewNo
Preview one calldry_run=true on the paid toolNo: the job is not submitted
Cap the callmax_spend_usdEnforced only when provided
Submitpaid tool with a new idempotency_keyYes

What a failed admission looks like

If Sume cannot reserve the estimated cost from the balance, the generation submit fails with 402 insufficient_credits before provider work starts. If the workspace has no remaining accepted capacity, the submit returns 429 queue_full. Treat these differently: a 402 means stop and ask for credit, a 429 means wait and retry with the same idempotency key.

The two argument sets

The Playbook C example in the docs shows the argument shape for a paid avatar create. This helper builds the preview and the submit from the same inputs, so the only difference is dry_run. It runs offline and prints both.

import json

def args(key: str, cap_usd: float, dry: bool) -> dict:
    if not key:
        raise ValueError("idempotency_key is required")
    return {
        "idempotency_key": key,
        "dry_run": dry,
        "max_spend_usd": cap_usd,
        "payload": {
            "avatar_handle": "studio_presenter",
            "input": {"type": "prompt",
                      "prompt": "A friendly studio presenter"},
        },
    }

if __name__ == "__main__":
    key = "avatar-create-2026-10-05-001"
    print(json.dumps(args(key, 2, True)))
    print(json.dumps(args(key, 2, False)))

Retries after a timeout

A preview is only half the safety story. The other half is what the agent does when a paid call times out. The jobs docs are direct: do not submit the original paid request again only because a local process timed out. Keep the idempotency_key, then read jobs_status or call jobs_wait on the job id you already have.

On remote MCP a wait slice is bounded: the default timeout_seconds is 50 and the cap is 55. When a slice expires, issue the wait again on the same ids. A 524 on jobs_wait is a transport failure, never a job outcome.

Agent rules

Put these in the agent's instructions or, better, in the code that wraps the tool.

  • Call balance_get first and stop if it cannot cover the preview.
  • Preview with dry_run=true, then read the estimate before you submit.
  • Use a new idempotency_key for each distinct job, and reuse it only for an exact retry.
  • After a timeout, do not submit again; poll with jobs_status or jobs_wait.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume