Kimi Code CLI YOLO and AFK modes with paid Sume tools over MCP
Kimi auto-approves MCP tool calls in YOLO and AFK modes. What that means for a Sume server with write access, and the limits to set first.

In Kimi Code CLI, YOLO and AFK modes grant MCP tool approvals automatically, so with a write-capable Sume connection the agent can submit paid generations without asking. Keep the connection read-only, or set a spend cap and a dry run habit, before you turn either mode on.
The Kimi page is direct about it: MCP tool calls prompt for confirmation, but in YOLO or AFK modes approvals are granted automatically, so enable them only for servers you fully trust.
What does Sume do when nobody approves a call?
Nothing extra. Hosted MCP has no human-approval step of its own. The tools and gates page is clear that idempotency_key is dedup rather than approval, that there is no mcp:paid scope, and that spend is decided by the wallet and admission.
Which controls still apply in auto-approve mode?
Three controls work without a human:
- Scope: an OAuth session with only
mcp:readcannot see or run paid tools at all. max_spend_usd: when the call includes it, the cap is enforced; it is not enforced when omitted.dry_run=true: returns a cost and admission preview and does not submit.- Wallet balance: a paid create the wallet cannot cover is refused.
How do you set up an unattended run safely?
For a run where nobody is watching, give Kimi a Read connection for research steps (crawl reads, catalog, job status), and a separate, scoped session for the one step that generates. Put the cap in the instruction: every paid call carries max_spend_usd, and a failed or timed-out step is retried with the same idempotency_key, never a new one.
Because the model can ignore a prompt rule, the strongest control is the credential: use a connection that cannot spend (OAuth Read) whenever the task does not need to.
| Control | Enforced by | Works unattended |
|---|---|---|
| OAuth Read only | Sume scope check | Yes |
max_spend_usd on the call | Sume, when provided | Only if the call includes it |
| Prompt rule to dry-run first | The model | No guarantee |
| Wallet balance | Sume admission | Yes |
What about long renders in an AFK session?
A wait that expires is not a failure. jobs_wait returns wait_slice_expired after at most 55 seconds, and the instruction is to wait again on the same ids, per the jobs page. Tell the agent that explicitly, or an unattended session may re-submit a paid create and pay twice.
What does the Kimi page not tell you?
It does not describe per-tool allowlists for MCP in those modes, so the all-or-nothing reading is the safe one: if the mode is on, assume every tool on every configured server is approved. That is why the credential matters more than any setting inside the CLI.
If you need both research and generation in one unattended run, consider two server entries with different credentials under different names, and enable only the one you need for a given task. The Sume side treats them as separate sessions, so the read-only one can never submit a paid job even if the model tries.
Sources
Related posts
More in Developers
- Kimi mcp test fails on Sume: URL, auth and scope checklist
If kimi mcp test fails or shows no Sume tools, check the endpoint URL, run kimi mcp auth, then separate a read-only grant from a broken connection.
- Kling motion control with avatar_id instead of image_url
Sume's Kling 3.0 Motion Control takes image_url or avatar_id/avatar_handle, never both. A ready avatar resolves server-side to its identity still.
- Kling motion control sync mode: a 30-second wait, then poll
Sume's sync and subscribe modes on Kling 3.0 Motion Control wait at most 30 seconds. A clip usually outlasts that, so poll status_url; do not resubmit.
- LangGraph 1.2 node timeout: the Sume video job keeps billing
LangGraph 1.2 adds run_timeout and idle_timeout per node. A timeout stops your node, not the Sume job it started: store the job id and re-poll.
Written by Sume