GPT-6 Astra on Sume holds a $10 reserve per turn: wallet check

Sume's billing code reserves 1000 cents on an Astra turn (2000 with Fast) and releases what the turn does not use. Check wallet headroom before you pick it.

4 min readSume
All posts

What the reserve is

When a Sume agent turn starts, the billing code reserves a ceiling against the wallet and, at the end, captures min(billable, reserved); unused reserve is released. For GPT-6 Astra that ceiling is 1000 cents, which is $10, and 2000 cents ($20) with the Fast checkbox. The repo comment explains it: Astra lists at $10 in and $50 out per million tokens, the same card as Fable 5 via OpenRouter, so it holds Fable's reserve rather than a rebuilt ceiling.

This is a hold, not a charge. A typical turn settles far below it. But a wallet with $6 free can fail to start an Astra turn that a Haiku turn would start without trouble.

Numbers behind it

The OpenAI model page, read 2026-10-08, lists the same short-context rates the Sume card uses. Fast is exactly twice Standard in the repo.

Astra rates and reserve (Sume agent model registry and rate cards on origin/main, read 2026-10-08; OpenAI page read 2026-10-08)
ItemStandardFast
Input per 1M tokens$10$20
Cached input per 1M tokens$1$2
Cache write per 1M tokens$12.50$25
Output per 1M tokens$50$100
Reserve held at turn start1000 cents ($10)2000 cents ($20)

What a turn actually costs

A 30,000-token prompt and a 2,000-token answer, with nothing cached, costs 30,000 x 10 / 1M = $0.30 plus 2,000 x 50 / 1M = $0.10, so $0.40. The hold is 25 times that. A long tool-heavy turn closer to 150,000 input and 3,000 output tokens is 1.50 + 0.15 = $1.65, still far under the ceiling, so the clip rarely bites.

The 1,050,000-token context and 128,000-token max output on the OpenAI page are the reasons the ceiling is not smaller: a single maximum-length answer at $50 per million output tokens is 128,000 x 50 / 1M = $6.40 by itself.

What to do

If you run unattended Format runs on Astra, keep at least $10 free in the wallet per concurrent run, and $20 if you use Fast. If you cannot, pick a model with a smaller reserve. The Format call docs say the receipt names the id that ran, so you can audit which runs held which ceiling.

How to read a failed start

Astra's reserve is a worst-case hold and it is released as unused. If a turn cannot start, the symptom is insufficient credit for the reserve, not a model error. The fix is a top-up or a smaller model, not a retry loop. Reserve and capture follow min(billable, reserved), so a turn can never be charged more than it held; the repo comment adds that a pathological turn clips rather than overdrafting.

Compare this with GPT-6 Sol at a $2 input and $10 output list, where the repo reserves 200 cents ($2), a tenth of Astra's hold at a fifth of the token price. The hold is proportional to the card but also rounded to a ceiling, which is why the ratio is not exact.

This is also why the hold matters more for teams that fan out many runs at once. Ten parallel Astra turns hold $100 together, even if each ends up costing under a dollar. A wallet with $50 in it would block the sixth start, even though the finished bill would have fit. Staggering the starts, or choosing Haiku 5.5 for the parallel planning and keeping Astra for a single final pass, keeps the holds small. The release of unused hold is why the balance recovers once the turns settle.

  • Standard Astra: hold 1000 cents. Fast: 2000 cents.
  • GPT-6 Sol and GPT-6.1 Sol: 200 cents each in the same table.

Sources

Related posts

More in Models

All Models posts

Written by Sume