GPT-6.1 Sol past 272K input tokens: the price doubles for the request

GPT-6.1 Sol charges 2x input and 1.5x output once a prompt passes 272K tokens. What that means for a long Format thread, with the vendor tables read Oct 3.

4 min readSume
All posts

Yes. OpenAI prices a GPT-6.1 Sol request at the long-context rate when its prompt goes over 272K input tokens: input and cache rates double and output rises 1.5x, and the higher rate covers the whole request, not only the tokens past the line. Below 272K you pay the standard card of $2 input and $10 output per 1M tokens.

The model's context window is 1,050,000 tokens, so a single request can sit well inside the window and still land on the expensive side of the line. For a video agent that keeps frames, transcripts and tool results in one conversation, that line is the number to watch.

What does OpenAI publish for each side of the line?

The pricing page lists two columns for GPT-6.1 Sol, short context and long context, and the model page states the 272K threshold in words. The figures below are copied from the Standard rows.

GPT-6.1 Sol Standard rates per 1M tokens, OpenAI pricing page, read 2026-10-03
Token typeShort context (up to 272K input)Long context (over 272K input)
Input$2$4
Cached input$0.10$0.20
Cache write$2.50$5
Output$10$15

What does one long request cost?

Take a request with 300,000 input tokens and 4,000 output tokens, nothing cached. At the long rate that is 0.3 x $4 = $1.20 for input plus 0.004 x $15 = $0.06 for output, so $1.26. The same 300,000 tokens at the short rate would have been $0.60 plus $0.04, or $0.64. The request is about twice as expensive because of where it landed, not because of what it did.

The same request with 270,000 input tokens stays on the short card: 0.27 x $2 = $0.54 plus $0.04 for output, $0.58. Thirty thousand more input tokens move the bill from $0.58 to about $1.26 in this example, which is why the threshold matters more than the average.

How does this meet a continued Format run?

A Sume Format run can continue an earlier run with previous_run_id, and the agent is replayed what the earlier run produced, so each continuation carries more history than the last. The runs doc describes that continuation, and the call doc says the model field picks the orchestrating LLM and the receipt echoes the id that ran.

Where a thread's per-turn prompt lands relative to 272K is not something a Format receipt reports as a tier. What the receipt does show is the money: usage.billable_amount_usd_micros is generation spend and excludes the LLM turn, while usage.debited_usd_micros is what the wallet deducted and includes the turn's own LLM row. Compare the two across turns of the same thread and the growth in the LLM share is visible.

What should a long thread do about it?

This post does not state how Sume bills a request above 272K; read your own receipts for that. The vendor side is the part that is fixed on the page today.

  • Start a fresh run with a short input when the old turns are not needed, instead of continuing a thread whose history you no longer use.
  • Pass the facts the next step needs in input rather than relying on a long replay.
  • Watch debited_usd_micros minus billable_amount_usd_micros per turn; a rising gap is the LLM cost of history.
  • Read the vendor tables again before budgeting: the 272K line and the multipliers are OpenAI's, and they can change.

Does the cache change the picture?

Cached input is cheaper on both sides of the line: $0.10 per 1M under 272K and $0.20 over it, against $2 and $4 for uncached input. A long thread that repeats a stable prefix benefits from caching, but the long-context multiplier still applies to the request as a whole. Caching lowers the rate for the repeated tokens; it does not move a 300K prompt back under the threshold.

OpenAI's pricing page also lists Batch and Flex rates that are half of Standard and a Fast mode at twice Standard, each with the same short and long split. Those are vendor service tiers; check what a given Sume run exposes before assuming they apply, and read the receipt rather than a list price when you need the actual spend.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume