OpenAI MCP has no per-call fee, but Sume generation still bills

OpenAI charges only tokens for MCP tool calls. A Sume generate_video or generate_image call still reserves wallet funds, so use dry_run and max_spend_usd.

4 min readSume
All posts

OpenAI says its remote MCP tool has no per-call fee: you pay for tokens, not for calling the tool. That does not make the work behind the tool free. A Sume generate_video or generate_image call reserves money from your Sume wallet when the job is accepted, so you have two bills: OpenAI tokens for the model loop and Sume USD for the media.

Two meters

Cost meters, read 2026-10-05
MeterCharged byWhat triggers it
MCP tool call itselfOpenAI: no per-call fee, tokens onlyTokens used to list tools, call them and read results
Media generationSume walletAccepted paid job; reservation at submit, capture on completion, release on failure
Rejected requestNobody402 insufficient_credits happens before provider work starts

Spend gates on the Sume side

Sume's hosted MCP gives you three controls. idempotency_key is required on write and paid tools, and it is for transport dedup, not human approval. dry_run=true returns an admission and cost preview without submitting. max_spend_usd is enforced only when you provide it. For bursts, the docs recommend generation_admission_preview or dry_run first. A normal single create does not need them.

{
  "idempotency_key": "clip-2026-10-05-001",
  "dry_run": true,
  "max_spend_usd": 2,
  "payload": {
    "prompt": "A ceramic mug rotating on a marble counter"
  }
}

How to use it

Call the paid tool once with dry_run true, read the estimate and balance, then call again with dry_run false and a fresh idempotency_key. Omit payload.model and Sume routes to sume/auto unless the user named a family. If the session shows insufficient_scope, the OAuth grant is read-only. Grant mcp:write at consent or use an API key.

On the OpenAI side, require_approval accepts always, never or filtered. Setting it to always on the paid tool gives a human a chance to see the dry-run numbers before real spend.

Do not forget

  • Retries with the same idempotency_key are the safe path. A different payload under the same key returns 409 idempotency_conflict.
  • queue_full and rate_limited are 429s, not billing events. A full queue means you stop submitting, not that money was spent.
  • The docs do not publish a total-cost formula across both vendors, so add the two meters yourself.

A budget in two parts

Put a number on each meter. For OpenAI, you estimate tokens from the tool list, the arguments and the results. For Sume, the preview is the number: a dry_run call returns the admission and cost estimate without creating a job. Set max_spend_usd to the amount you would accept for that one call. If the estimate is above it, the call should not go through.

An example: with max_spend_usd of 2 and a preview that estimates 2.40, do not submit. Either shorten the clip, lower the resolution, or ask a person to raise the cap. This turns an unbounded loop into a bounded one, which matters when an agent can call a paid tool many times.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume