Managed Agents bills $0.08 per session-hour; Sume's cap is separate
Claude Managed Agents charges tokens plus $0.08 per running session-hour. How that meter relates to Sume's max_spend_usd and spend cap.

Claude Managed Agents bills two things: tokens at the model rates, and session runtime at $0.08 per session-hour, counted only while the session status is running. If such an agent calls Sume's hosted MCP server, you have three meters running: Anthropic tokens, Anthropic runtime, and Sume's generation spend. Only the last is capped by Sume's max_spend_usd or, for Agent Completions, generation_spend_cap_usd.
Anthropic's meter
The pricing page says runtime is measured to the millisecond and accrues only while the session is running. Time spent idle, rescheduling or terminated does not count. Tokens are billed at the model prices, prompt-caching multipliers apply, and the Batch API discount does not, because sessions are stateful and interactive. Web search inside a session costs $10 per 1,000 searches.
| Meter | Rate | Notes |
|---|---|---|
| Session runtime | $0.08 per session-hour | Only while status is running |
| Tokens | Model rates | Cache multipliers apply |
| Batch discount | Not available | Sessions are interactive |
| Web search | $10 per 1,000 searches | Plus token costs |
Where the time goes with Sume
A Sume render can take minutes, and the agent waits with jobs_wait, which holds for up to 55 seconds per call on the remote server. While a tool call is in flight the session is working, so the wait is runtime. At $0.08 per hour, ten minutes of waiting is about one cent. The runtime meter is small next to a paid generation, which is the reason the Sume cap deserves the attention.
The Sume cap is a different kind of control. max_spend_usd is optional on hosted MCP and enforced only when you give it, so a managed agent should send it on every paid call. idempotency_key is required on paid and write calls. For unattended runs Sume's Agent Completions make generation_spend_cap_usd mandatory.
What to set
See auto permission mode and paid Sume calls in Managed Agents for the approval side.
- Set a per-call
max_spend_usdfrom a budget variable, not from model text. - Tell the agent to poll in slices and never to resubmit a create after a timeout.
- Track the two bills separately; a session that is cheap in runtime can still hold an expensive Sume job.
- Use a read-only OAuth scope for sessions that only need to look at jobs.
A budget sketch with published numbers
Anthropic's worked example prices a one-hour Opus 5 session with 50,000 input and 15,000 output tokens at $0.705 in total, of which $0.08 is runtime. Runtime was about 11 percent of that example. For a cheaper model the share would be higher, and for a session that waits on a long render with few tokens it can dominate the Anthropic side of the bill.
A Sume job is billed on top. Treat that as a fixed line you set with the cap, and treat the Anthropic side as the part that scales with how long the session stays running.
Sources
Related posts
More in Pricing
- Colossyan Professional: $59 for 30 minutes vs Sume per-minute rates
Colossyan lists Professional at $59 monthly or $30 annual for 30 NEO minutes and 3 seats. About $1.97 or $1.00 per video minute, against Sume at $11.04 to $33.
- One minute and one hour of AI narration: MAI-Voice vs Sume TTS cost
At 900 characters a minute, narration costs about 2 cents on MAI-Voice-2.1, 1.4 cents on Flash, 5 cents on Sume (job rounding). One hour: $1.19, $0.81, $2.58.
- Does Seedance aspect ratio change the price? 21:9 vs 9:16 on Sume
Aspect ratio barely moves Seedance prices on Sume: 9:16, 16:9 and 1:1 cost the same and 21:9 or 4:3 add under a percent. The 720p table and the pixel sizes.
- Omni Flash 1.1 price chart: 3 to 10 seconds at 4 resolutions
Every whole-second length from 3 to 10 seconds at 360p, 720p, 1080p and 4K for Gemini Omni Flash 1.1 on Sume, from the estimator that reserves your balance.
Written by Sume