402 automation_generation_spend_cap_exceeded in a scheduled run
The per-run generation cap rejected one job before it reserved credits. Only that job fails; the run is not canceled. Raise the cap or trim the plan.

402 automation_generation_spend_cap_exceeded means one generation job inside an automation run would have pushed that run past its generation spend cap, so the platform refused it before reserving any credits. Only that job fails; the automation run itself is not finalized or canceled, and the agent can carry on with cheaper work.
What the gate adds up
The gate sums the reserved plus the captured generation spend of the automation run and rejects the new job if adding its estimate would exceed the cap. Because it checks before the reserve, a rejected job leaves no hold behind.
The error body is a quota-category error with retryable: false and a fix_input next action. details carries four fields you can log: cap_usd_micros, current_usd_micros, requested_usd_micros and studio_agent_automation_run_id.
| Field | Meaning |
|---|---|
| cap_usd_micros | The cap in force for this run, in USD micros (1,000,000 = $1.00). |
| current_usd_micros | Reserved plus captured generation spend of the run so far. |
| requested_usd_micros | The estimate of the job that was refused. |
| studio_agent_automation_run_id | The automation run the job belonged to. |
Read the numbers
Compare current + requested with cap. If the sum is just over, the cap is too tight for the plan. If current is already near the cap, an earlier job in the same run used it up, and the run needs a smaller plan rather than a bigger cap.
The default cap for a scheduled run is $1.00 when none is set on the Action. A caller can only lower it for one run with generation_spend_cap_usd: the API clamps the request to the smaller of the request and the Action cap.
import os, requests
r = requests.post(
"https://api.sume.com/v1/actions/aut_REPLACE/runs",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]},
json={"input": {"topic": "weekly trends"},
"generation_spend_cap_usd": 0.5},
timeout=30,
)
print(r.status_code)Fixes in order
- Shrink the work: fewer clips, shorter durations, or a cheaper model in the instructions.
- Raise the Action's own cap in the dashboard if the plan is legitimately larger. A per-run request cannot raise it.
- Split a big plan across several runs so that each has its own ceiling.
An example reading
Say the log shows a cap of 1,000,000, a current of 800,000 and a requested of 350,000. The sum is 1,150,000, which is over by 150,000 micros, so the job is refused. Lowering the clip length in the schedule's instructions may bring the estimate under the line; raising the cap to 1,500,000 certainly will.
Log studio_agent_automation_run_id with the alert, so the person on call can find the run's receipt and see which earlier jobs consumed the budget.
Limits
This gate covers generation jobs only. The agent's own language-model turns bill a separate wallet and are not counted in these numbers. Interactive chats skip this gate entirely, so a plan that fails here can succeed when someone runs it by hand in chat.
The wire path above uses the /v1/actions namespace with an aut_ id; check the live API reference for the current route shape before you automate against it.
Sources
Related posts
More in Agents
- Agent Completion or three API calls for a render-trim-caption chain
If the steps are fixed, call the endpoints. If the task changes on every call, an Agent Completion with a required spend cap fits. How the two compare.
- Agent Completions input: JSON data the agent reads from a file
Send caller data in input, not in instruction. Sume writes it to /workspace/inputs/sume-action-input.json and tells the agent to read it as data. curl sample.
- Agent Completions request limits: 100,000 characters, 50 messages
The Sume Agent Completions schema caps instruction at 100,000 characters, messages at 50, and images at 30. Know where each limit is enforced before you send.
- Preflight a paid Sume MCP call: balance_get, then dry_run
Before a paid Sume tool call, an agent should read the balance and run the tool with dry_run=true. Then submit with a fresh idempotency_key and an optional cap.
Written by Sume