Copilot Studio hooks fail open: cap paid media spend server-side
Copilot Studio hooks can block a tool call, but a failed hook lets the agent continue. Put the media spend limit where it can't fail open: the Sume run cap.

Copilot Studio hooks can stop a paid tool call, but they are not a guarantee: Microsoft's page says a hook that fails, times out, or returns something unreadable lets the agent continue as if the hook returned nothing. So a hook is a good first gate for a paid media call, and a bad last one. The limit that cannot fail open is a spend cap enforced on the server that bills the generation, and Sume's Agent Completions, Format runs and schedules all carry one.
Everything below about hooks comes from Microsoft's Hooks (preview) page, read 2026-10-10. The Sume side comes from the docs pages linked in the sources. Nothing here says Copilot Studio can call Sume natively; it describes where to put each guard if an agent ends up with a tool that spends money.
What the hook page says it can and cannot do
A hook has an event and an action, and today the action is always a workflow. Unlike a tool, which the agent chooses, a hook runs every time its event fires. The page lists these events: start, user prompt submitted, error, pre tool use, post tool use, and after tool failure.
Only one of them can block anything. In Microsoft's words, pre tool use is the only event that can block an action; the others can add context or change values but cannot stop the agent. A pre tool use workflow receives the tool name and the parameters, and can return deny to block the call or modifiedParameters to rewrite it.
| Event | Can it block a call? | What it can change |
|---|---|---|
| Pre tool use | Yes, with a deny decision | Tool parameters, extra context |
| Post tool use | No | The result the model sees |
| After tool failure | No | Guidance passed with the failure |
| Error | No | Retry, skip or abort handling, and the user message |
| User prompt submitted | No | The prompt text, or hide the output |
The sentence that matters for spend
The same page says hooks do not stop the agent when they fail. If the workflow fails, times out, or returns something the agent cannot read, the agent continues as though the hook returned nothing. It also tells you not to rely on a hook as your only safeguard for a business-critical rule.
A rule such as "never spend more than $20 on one request" is business-critical. If the check lives in a workflow that can time out, a timeout equals permission. The page also says usage-based billing applies to agents powered by the GitHub Copilot harness, so the agent itself consumes credits; that is a separate meter from the generation your tool triggers elsewhere.
Where Sume enforces the limit
Sume puts the ceiling on the run, not on a caller-side check. The docs state it plainly for an unattended agent: an Agent Completion has tools and access to your generation wallet, so generation_spend_cap_usd is required, has no default, and a request without it fails with 400 invalid_request. In the chat UI an approval prompt protects you; a backend caller gets no prompt, and the cap replaces it.
Format runs and schedules work the same way with different defaults. A Format that never named a cap reports $400, a run can name its own number up to $500, and 0 is a 400. A schedule defaults to $1.00 per run and a caller can only lower that cap for a single run, never raise it. Receipts show the effective cap and the actual spend as usage.generation_spend_cap_usd_micros and usage.billable_amount_usd_micros.
| Surface | Cap field | If you send nothing |
|---|---|---|
| Agent Completions | generation_spend_cap_usd | 400 invalid_request, there is no default |
| Format run | generation_spend_cap_usd | The Format's cap; $400 when it never set one |
| Scheduled run | generation_spend_cap_usd | The schedule's cap; $1.00 when not set |
| Hosted MCP paid tool | max_spend_usd | No cap; Sume enforces it only when you provide it |
One honest gap, and two ways to close it
The last row matters. On hosted MCP, max_spend_usd is optional and Sume enforces it only when present. What gates MCP spend by default is the scope and the wallet: under OAuth mcp:read the paid tools are hidden, and write or paid calls need mcp:write or an API key plus an idempotency_key. If you expose Sume's hosted MCP to a Copilot Studio style agent, a pre tool use hook is the natural place to require max_spend_usd in the parameters and deny the call when it is missing.
The stronger design is to let the agent call a single capped endpoint instead. Give it a tool that starts an Agent Completion or a Format run from your backend, where your code sets generation_spend_cap_usd itself. The agent never chooses the number, and a hook failure cannot raise it.
- Use the hook for fast, friendly refusals: block calls with a missing or oversized
max_spend_usdbefore they reach Sume. - Use the server-side cap for the hard ceiling: it holds even if the hook workflow times out.
- Use a read-only OAuth session for exploration and add
mcp:writeonly for sessions that should spend. - Reuse a stable
idempotency_keyso a retried tool call does not submit a second paid job.
What the cap does not cover
Read the receipt carefully. Sume's billable_amount_usd_micros is the generation spend of that run, not the whole cost: it excludes the agent's own language-model turn, which bills a separate agent wallet, and GET /v1/usage remains the billing record. So a cap bounds media spend, while the model-token side of a Copilot Studio agent is governed by Microsoft's credits.
A minimal capped call from a backend looks like this, with the cap chosen by your code and not by the model:
curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: req-1042" \
-d '{"instruction":"Make one 9:16 product teaser from the attached photo.","generation_spend_cap_usd":5}'Sources
Related posts
More in Agents
- Gemini agent keeps memory for days; Sume API runs start fresh
Google's Gemini agent keeps one memory across devices and days. Sume's API runs start in a new thread, so state travels as input, files and previous_run_id.
- Guardrails for an AI agent that calls paid media APIs
Five guardrails for an agent that spends money: required cap, read before paying, idempotency keys, no secrets in logs, and a cancel path. Drawn from Sume docs.
- Planlock in front of Sume MCP: approve the plan, then the calls
Planlock is an MCP proxy that enforces a human-approved plan. Where it fits in front of Sume's hosted MCP, and which Sume gates still do work behind it.
- Porting chat-completions code to Sume Agent Completions
Agent Completions takes system and user messages, rejects assistant turns, returns a 202 receipt, and does not stream. Here is what to change when porting.
Written by Sume