Sume MCP dry run: next_step is only dry_run false, merge your payload
A paid Sume MCP dry run returns would_submit false and next_step arguments of just dry_run false. Re-send your original idempotency_key and payload with it.

When a paid Sume MCP tool runs with dry_run: true, the result is an mcp_paid_generation_dry_run object with would_submit: false, and agent.next_step points at the same tool with only { dry_run: false } in its arguments. That is a pointer, not a ready request: your caller must merge the original idempotency_key and payload before calling again.
Why a model gets this wrong
Agents that follow agent.next_step literally call the tool with arguments { "dry_run": false } and nothing else. The paid create then fails, because idempotency_key is required on write and paid tools and the payload is missing. The worse failure is a model that invents a fresh idempotency key for the real call: the dry run and the real submit then look unrelated, and a retry after a timeout cannot be deduplicated.
The server builds the dry-run next step without your inputs on purpose. The preview consumed the payload; the result does not echo it back as a request body.
The sequence
Keep one request object in your harness and flip a single flag.
| Step | dry_run | idempotency_key | Result |
|---|---|---|---|
| 1. Preview | true | K (your key) | mcp_paid_generation_dry_run, would_submit false, usage estimate |
| 2. Real call | false | same K | job accepted; wait with jobs_wait |
| 3. Retry after timeout | false | same K | same job, no second charge |
Harness code
Build the request once, then reuse it. This runs as written and prints the two argument sets.
import json, uuid
def build(payload: dict, key: str, dry_run: bool, cap: float) -> dict:
return {
'idempotency_key': key,
'max_spend_usd': cap,
'dry_run': dry_run,
'payload': payload,
}
key = 'launch-' + uuid.uuid4().hex[:12]
payload = {'prompt': 'A red kettle on a white table'}
preview = build(payload, key, True, 2.0)
real = build(payload, key, False, 2.0)
print(json.dumps(preview))
print(json.dumps(real))Limits
The dry run still runs the admission preview, and max_spend_usd is checked against it before the dry-run result returns, so a cap that is too low fails at step 1 with max_spend_exceeded rather than at submit (see the max_spend_usd exceeded post). A dry run is a price quote, not a reservation: balance can change before step 2.
Do not use dry runs for every single create. The server's own run contract tells agents to skip admission theater on ordinary single creates and small fan-outs, and to reserve previews for paid bursts or expensive packages.
Checklist before you ship
- Generate the idempotency_key once per intended job, before the dry run.
- Never let the model rewrite the payload between preview and submit.
- Keep max_spend_usd identical on both calls.
- On a timeout after the real call, resend with the same key instead of starting over.
Sources
Related posts
More in Developers
- Sume MCP generation_admission_rejected: read the preview first
generation_admission_rejected means the admission preview would not accept the request. Read the preview, then change the request or wait before retrying.
- Sume job_failed_terminal: retryable true means reuse the same key
On job_failed_terminal never resubmit the identical payload. If error.retryable is true, re-issue the create with the same idempotency_key; else fix the input.
- max_spend_usd on Sume MCP: floors to micros, missing estimate fails
max_spend_usd takes 0 to 10,000, floors to micros, and runs on the preview before dry_run returns. No usable estimate means missing_usage_estimate, not a pass.
- Sume MCP OAuth metadata: curl the well-known URLs before you connect
Before a client fails at sign-in, curl Sume's protected-resource and authorization-server metadata on mcp.sume.com and check the resource audience matches.
Written by Sume