Season budget ceiling: spend cap per episode times the queue size

Cap each Sume bulk item with generation_spend_cap_usd and the queue has a hard ceiling: cap times items, with concurrency limiting what is exposed at once.

5 min readSume
All posts

To bound what a season can cost on Sume, put generation_spend_cap_usd on every bulk item: the generation ceiling for the queue is the sum of those caps, and the most that can be exposed at one moment is about concurrency times the largest cap. The cap is your price control, so it is worth setting on purpose instead of inheriting the Format's default of $400.

Two facts from the Format call page set the arithmetic. A run's cap can be any number up to $500 and a number above the Format's own cap is honored, not clamped. And null runs at $500, which lifts the ceiling but does not remove it. 0 and anything above 500 are rejected with 400.

The worked ceiling

The caps below are illustrative, not prices. Replace them with the number you would be comfortable losing on one episode. Metered rates for each model are on Sume's API pricing page.

Queue ceiling for an illustrative $10 cap per episode (as of 2026-10-03)
Episodes in the queueConcurrencyCeiling for the queueMost exposed at once
82$80about $20
244$240about $40
10016$1,000about $160

What the cap does not include

  • usage.billable_amount_usd_micros is the generation spend attributed to the run. It excludes the agent's own LLM turn, so it is not the run's total cost and not an invoice.
  • For what the wallet actually deducted, read usage.debited_usd_micros; held_usd_micros and refunded_usd_micros show holds still open and holds given back.
  • GET /v1/usage stays the authoritative billing record, so reconcile the season against it.
  • The cap is not a wallet check. Wallet balance, admission and workspace concurrency still apply to every child, and a child that cannot start becomes a failed item while the queue keeps draining.

A pre-flight calculation

Do the sum before you submit, and refuse to send a queue whose ceiling exceeds your approved budget.

def queue_ceiling(caps_usd: list[float], concurrency: int) -> dict:
    if not caps_usd or not 1 <= len(caps_usd) <= 100:
        raise ValueError("a queue takes 1 to 100 items")
    if not 1 <= concurrency <= 16:
        raise ValueError("concurrency is 1 to 16")
    if any(c <= 0 or c > 500 for c in caps_usd):
        raise ValueError("each cap must be above 0 and at most 500")
    exposed = sum(sorted(caps_usd, reverse=True)[:concurrency])
    return {"ceiling": sum(caps_usd), "exposed_at_once": exposed}

print(queue_ceiling([10] * 24, 4))  # {'ceiling': 240, 'exposed_at_once': 40}

Reading the result

Each child's receipt carries the effective cap as usage.generation_spend_cap_usd_micros and the running spend as usage.billable_amount_usd_micros. A run that wants to spend past its cap lands as format_run_failed, so an episode that keeps failing is the first thing to check against its cap. Canceling a child with POST /v1/format-runs/{run_id}/cancel stops further spend, but generation already completed is billed.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume