automation_generation_spend_cap_exceeded: what a 402 cap error ends
The per-run cap rejects only the one generation that would cross it. current, requested and cap come back in micros so you can size the next request.

automation_generation_spend_cap_exceeded is a 402 raised before a generation is reserved, when the sum of what this automation run has already reserved and captured plus the new request would be over the run's cap. Only that one generation job fails. The OpenAPI description says the automation run itself is not finalized or cancelled, so your run can carry on with something cheaper or shorter.
What triggers it
The check applies only when the credential carries an automation run id. An interactive chat has none, so it skips the gate. Under a per-run lock the API sums reserved and captured generation spend for that run, adds the requested amount and compares the result with the cap. A run that was deliberately set up with no ceiling skips this ledger check, but wallet balance, generation admission and other limits still apply after it.
- It is
category: quota,stage: usage_reservation,retryable: false,next_action: fix_input. details.cap_usd_micros,current_usd_microsandrequested_usd_microsare in millionths of a USD.details.studio_agent_automation_run_idnames the run that hit its ceiling.
Read the headroom, then shrink the ask
The three amounts give you the headroom directly: cap minus current. If requested is bigger than that, the fix is a smaller request, not another attempt. The sample in the OpenAPI file has a cap of 1,000,000, 600,000 already spent and 500,000 requested, so the job is 100,000 over.
import json
BODY = '''{"error": {"code": "automation_generation_spend_cap_exceeded",
"details": {"cap_usd_micros": 1000000, "current_usd_micros": 600000,
"requested_usd_micros": 500000}}}'''
def headroom(body):
d = body["error"]["details"]
left = d["cap_usd_micros"] - d["current_usd_micros"]
over = d["requested_usd_micros"] - left
return left / 1e6, over / 1e6
left, over = headroom(json.loads(BODY))
print(f"left ${left:.2f}, request is ${over:.2f} over: cut it or raise the cap")Cap error or wallet error
Both are 402 and both are quota, so the status alone cannot tell them apart. insufficient_credits says add_funds. This one says the per-run ceiling is the limit, so topping up changes nothing. Cut the work, or change the cap on the run that owns it. For how Agent Completions needs a spend cap in the first place, see the related posts.
A fallback ladder for a capped run
Because only the one job fails, a run can degrade on purpose. Compute the headroom, keep the same prompt, and try a cheaper or shorter version. Stop when nothing fits, and log studio_agent_automation_run_id with the numbers so the person who owns the run can see why it stopped. Do not loop on the same body: the request is deterministic, so it returns the same refusal until the run's spend or its cap changes.
The Errors and credits page lists the general envelope, and the related posts on generation_spend_cap_usd cover how a cap is set on an Agent Completion.
Sources
Related posts
More in Developers
- Avatar batch: which limit hits first, writes per minute or the queue?
On Pro, 300 writes/min is far above the 24 accepted jobs (4 running, 20 queued). Queue capacity limits an avatar batch first, so submit in waves of that size.
- Avatar create 400: removed name and file fields and their replacements
Old avatar model-run requests that send name or file return 400 with details.fields listing replacements: avatar_handle and input.image_url. Fix both at once.
- 409 avatar_handle_reserved: why handles starting sume_ are off limits
Creating an avatar whose handle starts with sume_ returns 409 avatar_handle_reserved. What the error body says and how to pick a handle that passes.
- Avatar handle rules: 2 to 30 characters, with a Python check
A Sume avatar_handle allows a-z, 0-9, dot and underscore, 2 to 30 characters, no edge or doubled separators. A short Python check, plus the reserved prefix.
Written by Sume