Agent 365 cost management for Copilot Studio agents vs a per-run cap

Agent 365 sets spend policy per user group and adds Copilot Studio agents in October. A per-run cap on the API call bounds the one task. Know which you need.

4 min readSume
All posts

They answer different questions. Microsoft's Agent 365 cost management is a tenant-level control: policies about which AI models and capability levels different user groups can use, plus a consumption dashboard. A per-run cap, like Sume's generation_spend_cap_usd, is a call-level control for one task. If you run agents through Copilot Studio and also call outside APIs from them, you want both.

What Microsoft describes

The September post says Agent 365 cost management offers FinOps for AI, to manage usage-based spend, set spending guardrails, track costs and connect usage to value. It covers Copilot Cowork, WorkIQ, Code and Copilot Managed Runtime today, with agents built in Microsoft Copilot Studio scheduled for October. The dashboard shows usage trends, active users, credit consumption and estimated value. Microsoft says the capability is included with Microsoft cloud subscriptions.

Tenant policy versus call cap

Microsoft's post stresses that policies can restrict which models and capability levels groups may access, which is a good fit for keeping an expensive tier away from casual users. It does not claim to cap spend inside an external API called from an agent. Treat the dashboard as your view of Microsoft-side credits, and treat the API receipt as the source of truth for what the media provider charged.

Agent 365 cost management and a per-run API cap (read 2026-10-06)
QuestionAgent 365 cost managementSume per-run cap
Scope of controlGroups of users, models and capability levelsOne Agent Completion
When it appliesPolicy set ahead of use; dashboard afterBefore the run; no default, so the request fails without it
What it countsCredits for Microsoft agent workloadsGeneration spend on Sume
What it missesSpend inside third-party APIs an agent callsThe agent's own language-model turn, billed to the separate Agent wallet

The gap between the two

When a Copilot Studio agent calls a paid media API, Microsoft's dashboard shows the agent's credits and the API bills separately. Nothing joins the two unless you join them. Sume's receipt makes the join easy: every run has usage.billable_amount_usd_micros and generation_spend_cap_usd_micros, and the run id is stable as request_id in the signed webhook. Store that id next to the Microsoft-side record and you can reconcile later.

Because the Copilot Studio coverage is scheduled rather than shipped as of Microsoft's September post, confirm it is live in your tenant before you rely on it in a budget review.

What to set on the Sume side

On schedules, Sume's default cap is one dollar per run, and a per-run request value can only lower it. That keeps a cron mistake bounded even if nobody reads a dashboard.

  • generation_spend_cap_usd on every POST /v1/agent/completions, sized to one task.
  • An Idempotency-Key so retries return the original receipt with idempotency_hit: true.
  • A webhook with outcome branching, so a degraded run (billed, no structured output) opens a review rather than a silent retry.

Sources

More in Pricing

All Pricing posts

Written by Sume