Agent retry budget for paid video calls: stop on 402
An agent that calls a paid video API needs three counters: submits, estimated dollars and consecutive 402s. The stop rule, and how Sume's spend cap backs it up.

Give an agent that calls a paid video API three hard limits in its own code: a maximum number of submits, a maximum estimated spend, and a rule that the first 402 stops the loop. Do not let the model decide whether to try again. A language model asked to continue after a payment error will often continue, and each continuation on Sume is a new reserved hold against a wallet only a human can top up.
Why 402 is a stop, not a retry
On Sume, 402 insufficient_credits carries next_action: add_funds, and the credits docs describe top-ups as a dashboard action, with no public top-up endpoint an agent could call. A 4xx at create means nothing ran and nothing was charged, so stopping is free; looping is only noise against the write budget. The same holds for the Format-level cap: a run that would spend past its generation_spend_cap_usd ends as format_run_failed, and re-sending the same run into the same cap fails again.
Contrast this with transient classes. 429 rate_limited, queue_full and 503 provider_capacity_exceeded are worth retrying with backoff and the same Idempotency-Key. The agent loop needs to separate those two families in code, not in the prompt.
Three counters in 20 lines
The class below tracks remaining submits and remaining estimated dollars, refuses an action that would cross either, and treats 402 as terminal. It is pure Python and runs as-is; the estimates would come from the catalog or your own price table. Note that a retry on 503 still spends a submit: the budget counts attempts, not successes, which is what stops a flapping provider from draining the loop.
class Budget:
def __init__(self, max_submits=3, max_usd=6.0):
self.left_submits, self.left_usd = max_submits, max_usd
def allow(self, est_usd):
return self.left_submits > 0 and est_usd <= self.left_usd
def spend(self, est_usd):
self.left_submits -= 1
self.left_usd -= est_usd
def step(budget, est_usd, status):
if not budget.allow(est_usd):
return "stop: budget"
if status == 402:
return "stop: ask a human to add funds"
budget.spend(est_usd)
return "ok" if status == 202 else f"retry later ({status})"
b = Budget()
for status in (202, 503, 202, 202, 402):
print(step(b, 2.0, status), b.left_submits, b.left_usd)Let the platform back you up
Client counters are the first line. The second is the run's own cap: omit generation_spend_cap_usd and the Format cap applies ($400 by default), pass a number up to 500 and it is honoured even above the Format cap, pass null and you get $500, not unlimited. Zero or more than 500 is a 400. For an unattended agent, set the number explicitly to what one task is worth.
After the fact, read spend from /v1/usage?run_id= summaries rather than adding rows, since reserved, captured and refunded rows would double count. Compare usage.billable_amount_usd_micros to the cap; billable excludes the agent's own LLM turn.
| Signal | Agent action | Note |
|---|---|---|
| 402 insufficient_credits | Stop; hand to a human | Dashboard top-up only |
| format_run_failed at the cap | Stop; raise the cap deliberately | Same cap fails again |
| 429 rate_limited | Back off with retry-after | Same key |
| 429 queue_full | Wait for a job to finish | Same key |
| 503 provider_capacity_exceeded | Retry later | Counts against the submit budget |
Hand-off text
When the loop stops on 402, have the agent output the facts a person needs: the run or job id if one existed, the estimated cost, the request id from x-sume-request-id, and the sentence that credits must be added in the dashboard. Then end the task. Do not poll the balance in a tight loop; one read after the human replies is enough.
Reading the counters back
After the loop ends, log the three counters with the run id and the request id. Over a week they show whether the agent is bumping into the money limit, the retry limit or the submit limit, and each points at a different change: a higher wallet, a better backoff, or a smaller task. Keep the thresholds in configuration, not in the prompt, so a person can raise them deliberately. If the agent is allowed to call tools through a hosted endpoint, remember that each tool call spends the write budget once and status polls spend none of it, so counting submits and counting tool calls are not the same number.
Sources
Related posts
More in Agents
- Claude Code 2.1.288 background session fix: Sume jobs keep running
Claude Code 2.1.288 fixed background sessions ending on plugin reload. A Sume job outlives the session; how to find it and avoid paying twice.
- Claude Code mcp_tool hook skipped on SessionStart: Sume balance check
Claude Code's mcp_tool hooks are skipped on SessionStart and Setup because no MCP client exists yet. Run a Sume balance check at PreToolUse instead.
- Claude Code plugin agents honor disallowedTools: block Sume paid tools
Since 2.1.288 plugin-defined agents run with their own prompt, tools, disallowedTools and effort. A reviewer agent that can read Sume jobs but never create one.
- Claude Code RETRY_WATCHDOG gives up after 3 timeouts: Sume jobs
Unattended Claude Code sessions using CLAUDE_CODE_RETRY_WATCHDOG now stop after three timeouts. Persist Sume job ids so a restart resumes, not resubmits.
Written by Sume