GPT-6 Astra on Sume holds a $10 reserve per turn: wallet check
Sume's billing code reserves 1000 cents on an Astra turn (2000 with Fast) and releases what the turn does not use. Check wallet headroom before you pick it.

What the reserve is
When a Sume agent turn starts, the billing code reserves a ceiling against the wallet and, at the end, captures min(billable, reserved); unused reserve is released. For GPT-6 Astra that ceiling is 1000 cents, which is $10, and 2000 cents ($20) with the Fast checkbox. The repo comment explains it: Astra lists at $10 in and $50 out per million tokens, the same card as Fable 5 via OpenRouter, so it holds Fable's reserve rather than a rebuilt ceiling.
This is a hold, not a charge. A typical turn settles far below it. But a wallet with $6 free can fail to start an Astra turn that a Haiku turn would start without trouble.
Numbers behind it
The OpenAI model page, read 2026-10-08, lists the same short-context rates the Sume card uses. Fast is exactly twice Standard in the repo.
| Item | Standard | Fast |
|---|---|---|
| Input per 1M tokens | $10 | $20 |
| Cached input per 1M tokens | $1 | $2 |
| Cache write per 1M tokens | $12.50 | $25 |
| Output per 1M tokens | $50 | $100 |
| Reserve held at turn start | 1000 cents ($10) | 2000 cents ($20) |
What a turn actually costs
A 30,000-token prompt and a 2,000-token answer, with nothing cached, costs 30,000 x 10 / 1M = $0.30 plus 2,000 x 50 / 1M = $0.10, so $0.40. The hold is 25 times that. A long tool-heavy turn closer to 150,000 input and 3,000 output tokens is 1.50 + 0.15 = $1.65, still far under the ceiling, so the clip rarely bites.
The 1,050,000-token context and 128,000-token max output on the OpenAI page are the reasons the ceiling is not smaller: a single maximum-length answer at $50 per million output tokens is 128,000 x 50 / 1M = $6.40 by itself.
What to do
If you run unattended Format runs on Astra, keep at least $10 free in the wallet per concurrent run, and $20 if you use Fast. If you cannot, pick a model with a smaller reserve. The Format call docs say the receipt names the id that ran, so you can audit which runs held which ceiling.
How to read a failed start
Astra's reserve is a worst-case hold and it is released as unused. If a turn cannot start, the symptom is insufficient credit for the reserve, not a model error. The fix is a top-up or a smaller model, not a retry loop. Reserve and capture follow min(billable, reserved), so a turn can never be charged more than it held; the repo comment adds that a pathological turn clips rather than overdrafting.
Compare this with GPT-6 Sol at a $2 input and $10 output list, where the repo reserves 200 cents ($2), a tenth of Astra's hold at a fifth of the token price. The hold is proportional to the card but also rounded to a ceiling, which is why the ratio is not exact.
This is also why the hold matters more for teams that fan out many runs at once. Ten parallel Astra turns hold $100 together, even if each ends up costing under a dollar. A wallet with $50 in it would block the sixth start, even though the finished bill would have fit. Staggering the starts, or choosing Haiku 5.5 for the parallel planning and keeping Astra for a single final pass, keeps the holds small. The release of unused hold is why the balance recovers once the turns settle.
- Standard Astra: hold 1000 cents. Fast: 2000 cents.
- GPT-6 Sol and GPT-6.1 Sol: 200 cents each in the same table.
Sources
Related posts
More in Models
- GPT Image 2.5 draft grid: four low-quality images for 4 cents
Four GPT Image 2.5 drafts in one Sume call at quality low cost 4 cents; a dollar buys 25 such calls. The arithmetic, the n field and when to go to high.
- Grok Imagine on Sume: 3 cents, one image per call, 9:20 phone ratios
Grok Imagine bills 3 cents per image on Sume, returns one image per call, lists 13 ratios including 9:20 and 19.5:9 phone shapes, and supports edits.
- H3 on vLLM-Omni: 87 s on 4 B300 for one clip, and the cost
MiniMax says an 8.7 s H3 clip takes about 87 s on 4 B300 GPUs. Turn that into GPU-seconds, a break-even rate, and compare with a 9 s Sume job.
- Haiku 4.5 in Sume's registry now points to Haiku 5.5: $1 to $0.10
Sume's registry retired Haiku 4.5 onto Haiku 5.5 with the successor set to always. List input drops from $1 to $0.10 per million; old rows keep the 4.5 card.
Written by Sume