Comfy Agent Auto mode burns credits per token: how to stay capped
Comfy Agent bills tokens from the same Comfy Credits pool and needs a positive balance. Auto mode skips run approvals. What to cap on Sume instead.

Comfy Agent charges tokens from the same Comfy Credits pool you already use for generations, and you need a positive balance to keep running it, though not necessarily a subscription. In Auto mode it runs workflows without asking for each run. The page I read documents a Stop control for the current response and the balance requirement, and I found no per-chat spending limit on it, so the credit balance is your ceiling. On Sume, the matching ceiling is a number you set on each run.
Rates and modes are from Comfy's Agent documentation, read on 2026-10-03; Sume's side is in Agent Completions and MCP tools and gates.
What does Comfy charge?
The page lists per-token rates for the model it names. These are credits per 1,000 tokens.
| Token type | Credits per 1K tokens |
|---|---|
| Input | 1.055 |
| Output | 5.275 |
| 5-minute cache write | 1.319 |
| 1-hour cache write | 2.110 |
| Cache hit | 0.106 |
What do Ask and Auto change?
Ask requests approval before each workflow run; Auto runs workflows without asking each time. Token cost accrues while the Agent plans and iterates, so Auto mode combines two meters: tokens for thinking and credits for each generation it launches. Select Stop in the message box to end the current response.
How do I cap the Sume side when Comfy Agent calls out?
If a Comfy workflow reaches Sume, for example through a custom node calling the API, the credit pool above does not apply. Sume has its own balance and admission. Send generation_spend_cap_usd on every Agent Completion, or max_spend_usd on every paid MCP call, because Sume enforces the MCP cap only when you send it. Use dry_run=true to preview before a burst, and generation_admission_preview before a wave.
A balance shortfall fails at submit with 402 insufficient_credits before provider work starts, so an agent loop can be told to stop on that code instead of retrying.
What is a safe default?
One practical habit: check the balance before leaving an Auto session unattended. Because cache hits are priced far below fresh input at 0.106 credits per 1K tokens against 1.055, long sessions that reuse context cost less per token than many short ones that rebuild it, but the page still gives no total limit, so the balance is the only brake.
- Start in Ask mode on a new workflow, switch to Auto only after one clean run.
- Keep a small positive balance on the account you hand to an unattended agent.
- Cap each Sume call in dollars, so Comfy's token meter and Sume's generation meter are bounded separately.
- Stop on
402, and readretry-afteron429rather than looping.
Sources
Related posts
More in Pricing
- Copilot Studio computer use costs 5 credits a step: set a ceiling
Computer use bills 5 Copilot Credits per step, 15 on a premium model. The step count is unknown up front. A Sume run takes a dollar cap before it starts.
- How much do 1,000 AI images cost on the Sume image API?
Per-image list prices from the Sume image catalog for 18 model rows, multiplied out to 1,000 images, with the rows where the listed price is only a default.
- Cost of 12 avatar clips (4 each at 15, 30, 45 s) on Sume
Twelve talking-avatar clips, 360 seconds total, cost $66.24 on Standard, $88.20 on Plus and $198.00 on Max at Sume's listed rates, plus $0.95 once.
- Cost per row of a Sume bulk run: add up debited_usd_micros
Read usage.debited_usd_micros on each child receipt, wait for final to be true, and treat null as unknown. Why billable_amount alone understates a bulk run.
Written by Sume