Claude Code PreToolUse hook: block paid Sume calls without a spend cap
A PreToolUse hook that denies Sume generate and TTS tools when max_spend_usd or idempotency_key is missing. Python hook with the settings.json matcher, tested.

To stop a coding agent from running a paid Sume tool without a spend cap, add a Claude Code PreToolUse hook that matches the Sume hosted MCP tools and denies any call whose input has no max_spend_usd or no idempotency_key. Sume's own rule is that paid and write tools need an idempotency_key, with dry_run and max_spend_usd as optional guards, per Sume hosted MCP tools and gates. A hook turns the optional guard into a house rule, and it runs before the call leaves your machine, so the check does not depend on the model remembering a prompt instruction.
The hook below is local policy, not a Sume feature. It is a small script that reads one JSON object on stdin and either stays silent or prints a deny decision.
How does Claude Code call the hook?
According to the Claude Code hooks reference, a PreToolUse entry in settings has a matcher and a list of command hooks; MCP tools are named mcp__<server>__<tool>; the hook receives tool_name and tool_input as JSON on stdin; and a deny is either exit code 2, or exit 0 with hookSpecificOutput carrying permissionDecision: "deny" and a permissionDecisionReason that Claude sees. If you add the server as sume, the tools are mcp__sume__generate_video, mcp__sume__tts_create and so on.
What is the hook script?
import json, sys
PAID = ("generate_image", "generate_video", "tts_create", "music_create")
def decide(event):
name = event.get("tool_name", "")
if not name.startswith("mcp__sume__") or name.split("__")[-1] not in PAID:
return None
args = event.get("tool_input") or {}
missing = [k for k in ("max_spend_usd", "idempotency_key") if not args.get(k)]
if not missing:
return None
return {"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "deny",
"permissionDecisionReason": "Add " + " and ".join(missing) + " to this Sume call.",
}}
ok = {"tool_name": "mcp__sume__generate_video",
"tool_input": {"max_spend_usd": 2, "idempotency_key": "k1", "prompt": "x"}}
assert decide(ok) is None
bad = {"tool_name": "mcp__sume__tts_create", "tool_input": {"idempotency_key": "k"}}
assert "max_spend_usd" in decide(bad)["hookSpecificOutput"]["permissionDecisionReason"]
if __name__ == "__main__":
raw = sys.stdin.read().strip()
out = decide(json.loads(raw)) if raw else None
if out:
print(json.dumps(out))
print("hook ok", file=sys.stderr)Where does it get wired in?
Wire it in .claude/settings.json with the matcher mcp__sume__.* and the command pointing at the script, as in the reference's own examples. Reads such as jobs_status, jobs_result and tools_list pass untouched because the script only inspects the four creating tools, and a missing hook script should be treated as an outage, not a bypass, so keep it in the repository next to the settings file.
Test it by hand before trusting it: pipe one JSON line into the script and read the output, and pipe an empty string to confirm it stays silent, since a hook that crashes on empty input can block every tool call in the session. A deny is not a dead end. Claude receives the reason text and usually retries the call with the field filled in, which is the behaviour you want: the spend cap becomes part of how the agent asks, and a human sees it in the transcript.
| Guard | Who enforces it | Stops |
|---|---|---|
| idempotency_key | Sume (required on paid and write tools) | A repeated call billing twice |
| max_spend_usd | Your hook (optional in Sume) | A single call above your cap |
| dry_run | Your hook or prompt (optional in Sume) | Spend before an estimate is read |
| OAuth without mcp:write | Sume | Any write or paid tool |
Should the hook also check the amount?
Keep the cap realistic. A hook that blocks every call gets disabled by the first person it annoys, so set the required field and let your cap value live in the agent's instructions. This script checks presence, not size, on purpose: Sume does not publish a spend schedule in this post's sources, and inventing a number here would be a guess. Treat the cap as a per-call ceiling that your team agrees on, write it in the project instructions, and review the transcripts occasionally to see whether agents are filling the field with sensible values or with a token that merely satisfies the check. If you want a size ceiling, read the value as a number in decide and compare it to a limit you own.
Pair the hook with the two guards Sume already documents. dry_run=true is an optional admission and cost preview that does not submit the job, and the docs recommend a preview, either dry_run or generation_admission_preview, before expensive bursts, while a normal single create does not need one. And an OAuth session with only mcp:read cannot call write or paid tools at all: they return insufficient_scope. A hook is the right layer for team policy, but scope is the stronger one, so sign agents in read-only until a person has watched the first paid call succeed.
Sources
Related posts
More in Developers
- AGENTS.md rules for a coding agent that calls Sume over hosted MCP
Six AGENTS.md lines that stop a coding agent double-billing Sume video jobs: idempotency_key, jobs_wait slices, dry_run, plus a CI lint script.
- curl and jq script to test a new Sume image model id in one command
A 6-line shell script that posts one prompt to Sume POST /v1/images for any model id and prints the url, cost and status, to vet a gpt-image-1 replacement fast.
- Cursor mcp.json for the hosted Sume server: env interpolation or OAuth
Cursor reads .cursor/mcp.json with url and headers and supports ${env:NAME}. Pass the Sume key from the environment, or omit headers and use OAuth. Validated.
- Django webhook view for an AI video job: csrf_exempt and HMAC check
A Django view that verifies Sume's x-sume-webhook-signature over timestamp.raw_body, rejects an empty secret, and handles job.completed for a /v1/videos clip.
Written by Sume