Agent retry budget for paid video calls: stop on 402

An agent that calls a paid video API needs three counters: submits, estimated dollars and consecutive 402s. The stop rule, and how Sume's spend cap backs it up.

6 min readSume
All posts

Give an agent that calls a paid video API three hard limits in its own code: a maximum number of submits, a maximum estimated spend, and a rule that the first 402 stops the loop. Do not let the model decide whether to try again. A language model asked to continue after a payment error will often continue, and each continuation on Sume is a new reserved hold against a wallet only a human can top up.

Why 402 is a stop, not a retry

On Sume, 402 insufficient_credits carries next_action: add_funds, and the credits docs describe top-ups as a dashboard action, with no public top-up endpoint an agent could call. A 4xx at create means nothing ran and nothing was charged, so stopping is free; looping is only noise against the write budget. The same holds for the Format-level cap: a run that would spend past its generation_spend_cap_usd ends as format_run_failed, and re-sending the same run into the same cap fails again.

Contrast this with transient classes. 429 rate_limited, queue_full and 503 provider_capacity_exceeded are worth retrying with backoff and the same Idempotency-Key. The agent loop needs to separate those two families in code, not in the prompt.

Three counters in 20 lines

The class below tracks remaining submits and remaining estimated dollars, refuses an action that would cross either, and treats 402 as terminal. It is pure Python and runs as-is; the estimates would come from the catalog or your own price table. Note that a retry on 503 still spends a submit: the budget counts attempts, not successes, which is what stops a flapping provider from draining the loop.

class Budget:
    def __init__(self, max_submits=3, max_usd=6.0):
        self.left_submits, self.left_usd = max_submits, max_usd

    def allow(self, est_usd):
        return self.left_submits > 0 and est_usd <= self.left_usd

    def spend(self, est_usd):
        self.left_submits -= 1
        self.left_usd -= est_usd

def step(budget, est_usd, status):
    if not budget.allow(est_usd):
        return "stop: budget"
    if status == 402:
        return "stop: ask a human to add funds"
    budget.spend(est_usd)
    return "ok" if status == 202 else f"retry later ({status})"

b = Budget()
for status in (202, 503, 202, 202, 402):
    print(step(b, 2.0, status), b.left_submits, b.left_usd)

Let the platform back you up

Client counters are the first line. The second is the run's own cap: omit generation_spend_cap_usd and the Format cap applies ($400 by default), pass a number up to 500 and it is honoured even above the Format cap, pass null and you get $500, not unlimited. Zero or more than 500 is a 400. For an unattended agent, set the number explicitly to what one task is worth.

After the fact, read spend from /v1/usage?run_id= summaries rather than adding rows, since reserved, captured and refunded rows would double count. Compare usage.billable_amount_usd_micros to the cap; billable excludes the agent's own LLM turn.

read 2026-10-03
SignalAgent actionNote
402 insufficient_creditsStop; hand to a humanDashboard top-up only
format_run_failed at the capStop; raise the cap deliberatelySame cap fails again
429 rate_limitedBack off with retry-afterSame key
429 queue_fullWait for a job to finishSame key
503 provider_capacity_exceededRetry laterCounts against the submit budget

Hand-off text

When the loop stops on 402, have the agent output the facts a person needs: the run or job id if one existed, the estimated cost, the request id from x-sume-request-id, and the sentence that credits must be added in the dashboard. Then end the task. Do not poll the balance in a tight loop; one read after the human replies is enough.

Reading the counters back

After the loop ends, log the three counters with the run id and the request id. Over a week they show whether the agent is bumping into the money limit, the retry limit or the submit limit, and each points at a different change: a higher wallet, a better backoff, or a smaller task. Keep the thresholds in configuration, not in the prompt, so a person can raise them deliberately. If the agent is allowed to call tools through a hosted endpoint, remember that each tool call spends the write budget once and status polls spend none of it, so counting submits and counting tool calls are not the same number.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume