GPT-6.1 Sol past 272K input tokens: the price doubles for the request
GPT-6.1 Sol charges 2x input and 1.5x output once a prompt passes 272K tokens. What that means for a long Format thread, with the vendor tables read Oct 3.

Yes. OpenAI prices a GPT-6.1 Sol request at the long-context rate when its prompt goes over 272K input tokens: input and cache rates double and output rises 1.5x, and the higher rate covers the whole request, not only the tokens past the line. Below 272K you pay the standard card of $2 input and $10 output per 1M tokens.
The model's context window is 1,050,000 tokens, so a single request can sit well inside the window and still land on the expensive side of the line. For a video agent that keeps frames, transcripts and tool results in one conversation, that line is the number to watch.
What does OpenAI publish for each side of the line?
The pricing page lists two columns for GPT-6.1 Sol, short context and long context, and the model page states the 272K threshold in words. The figures below are copied from the Standard rows.
| Token type | Short context (up to 272K input) | Long context (over 272K input) |
|---|---|---|
| Input | $2 | $4 |
| Cached input | $0.10 | $0.20 |
| Cache write | $2.50 | $5 |
| Output | $10 | $15 |
What does one long request cost?
Take a request with 300,000 input tokens and 4,000 output tokens, nothing cached. At the long rate that is 0.3 x $4 = $1.20 for input plus 0.004 x $15 = $0.06 for output, so $1.26. The same 300,000 tokens at the short rate would have been $0.60 plus $0.04, or $0.64. The request is about twice as expensive because of where it landed, not because of what it did.
The same request with 270,000 input tokens stays on the short card: 0.27 x $2 = $0.54 plus $0.04 for output, $0.58. Thirty thousand more input tokens move the bill from $0.58 to about $1.26 in this example, which is why the threshold matters more than the average.
How does this meet a continued Format run?
A Sume Format run can continue an earlier run with previous_run_id, and the agent is replayed what the earlier run produced, so each continuation carries more history than the last. The runs doc describes that continuation, and the call doc says the model field picks the orchestrating LLM and the receipt echoes the id that ran.
Where a thread's per-turn prompt lands relative to 272K is not something a Format receipt reports as a tier. What the receipt does show is the money: usage.billable_amount_usd_micros is generation spend and excludes the LLM turn, while usage.debited_usd_micros is what the wallet deducted and includes the turn's own LLM row. Compare the two across turns of the same thread and the growth in the LLM share is visible.
What should a long thread do about it?
This post does not state how Sume bills a request above 272K; read your own receipts for that. The vendor side is the part that is fixed on the page today.
- Start a fresh run with a short input when the old turns are not needed, instead of continuing a thread whose history you no longer use.
- Pass the facts the next step needs in
inputrather than relying on a long replay. - Watch
debited_usd_microsminusbillable_amount_usd_microsper turn; a rising gap is the LLM cost of history. - Read the vendor tables again before budgeting: the 272K line and the multipliers are OpenAI's, and they can change.
Does the cache change the picture?
Cached input is cheaper on both sides of the line: $0.10 per 1M under 272K and $0.20 over it, against $2 and $4 for uncached input. A long thread that repeats a stable prefix benefits from caching, but the long-context multiplier still applies to the request as a whole. Caching lowers the rate for the repeated tokens; it does not move a 300K prompt back under the threshold.
OpenAI's pricing page also lists Batch and Flex rates that are half of Standard and a Fast mode at twice Standard, each with the same short and long split. Those are vendor service tiers; check what a given Sume run exposes before assuming they apply, and read the receipt rather than a list price when you need the actual spend.
Sources
Related posts
More in Pricing
- GPT Image 2.5 portrait 1024x1536 costs less than square
On fal, GPT Image 2.5 high is $0.04116 at 1024x1536 and $0.05268 at 1024x1024, though portrait has more pixels. Sume at x 1.25: about $0.051 vs $0.066.
- GPT Image 2.5 price on Sume: 1K, 1080p, 2K and 4K at low, medium, high
GPT Image 2.5 on Sume at high quality: about $0.066 at 1024x1024, $0.050 at 1920x1080, $0.069 at 2560x1440, $0.125 at 4K. Low is under a cent to 1.4 cents.
- Grok Imagine Video 1.5 on Sume: one per-second rate for 480p and 720p
Sume's catalog prices grok-imagine-video-1.5 at one per-second list rate for 480p and 720p. Read the live rate from the catalog before budgeting.
- H3 Max Recast 768p or 1080p: draft cheap, publish in HD
Sume bills H3 Max Recast at $0.375 a second at 768p and $0.5625 at 1080p. Budget a 768p test pass and one 1080p final, and what each plan costs.
Written by Sume