automation_generation_spend_cap_exceeded: what a 402 cap error ends

The per-run cap rejects only the one generation that would cross it. current, requested and cap come back in micros so you can size the next request.

4 min readSume
All posts

automation_generation_spend_cap_exceeded is a 402 raised before a generation is reserved, when the sum of what this automation run has already reserved and captured plus the new request would be over the run's cap. Only that one generation job fails. The OpenAPI description says the automation run itself is not finalized or cancelled, so your run can carry on with something cheaper or shorter.

What triggers it

The check applies only when the credential carries an automation run id. An interactive chat has none, so it skips the gate. Under a per-run lock the API sums reserved and captured generation spend for that run, adds the requested amount and compares the result with the cap. A run that was deliberately set up with no ceiling skips this ledger check, but wallet balance, generation admission and other limits still apply after it.

  • It is category: quota, stage: usage_reservation, retryable: false, next_action: fix_input.
  • details.cap_usd_micros, current_usd_micros and requested_usd_micros are in millionths of a USD.
  • details.studio_agent_automation_run_id names the run that hit its ceiling.

Read the headroom, then shrink the ask

The three amounts give you the headroom directly: cap minus current. If requested is bigger than that, the fix is a smaller request, not another attempt. The sample in the OpenAPI file has a cap of 1,000,000, 600,000 already spent and 500,000 requested, so the job is 100,000 over.

import json

BODY = '''{"error": {"code": "automation_generation_spend_cap_exceeded",
 "details": {"cap_usd_micros": 1000000, "current_usd_micros": 600000,
 "requested_usd_micros": 500000}}}'''

def headroom(body):
    d = body["error"]["details"]
    left = d["cap_usd_micros"] - d["current_usd_micros"]
    over = d["requested_usd_micros"] - left
    return left / 1e6, over / 1e6

left, over = headroom(json.loads(BODY))
print(f"left ${left:.2f}, request is ${over:.2f} over: cut it or raise the cap")

Cap error or wallet error

Both are 402 and both are quota, so the status alone cannot tell them apart. insufficient_credits says add_funds. This one says the per-run ceiling is the limit, so topping up changes nothing. Cut the work, or change the cap on the run that owns it. For how Agent Completions needs a spend cap in the first place, see the related posts.

A fallback ladder for a capped run

Because only the one job fails, a run can degrade on purpose. Compute the headroom, keep the same prompt, and try a cheaper or shorter version. Stop when nothing fits, and log studio_agent_automation_run_id with the numbers so the person who owns the run can see why it stopped. Do not loop on the same body: the request is deterministic, so it returns the same refusal until the run's spend or its cap changes.

The Errors and credits page lists the general envelope, and the related posts on generation_spend_cap_usd cover how a cap is set on an Agent Completion.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume