GPT-6.1 Sol cache write is $2.50 per 1M: when does it pay off?

GPT-6.1 Sol lists $2 input, $2.50 cache write and $0.10 cached input per 1M tokens. One reuse of a prefix already beats sending it fresh; the math.

5 min readSume
All posts

A cache write on GPT-6.1 Sol costs $2.50 per 1M tokens, against $2.00 for the same tokens sent as plain input, so writing is 25 percent dearer than not caching. Reading the prefix back costs $0.10 per 1M. That means a prefix you reuse even once is cheaper cached: write plus one read is $2.60 per 1M, while sending it fresh twice is $4.00.

The numbers come from the OpenAI API changelog, entry of Sep 29: GPT-6.1 Sol at $2 input, $0.10 cached input, $2.50 cache write and $10 output per 1M tokens, up to 272K. The same page lists GPT-6 Sol (Sep 22) at $2 input, $0.20 cached input and $10 output, with no cache-write price shown. That page does not say when a write is charged or how long an entry lives, so this post prices the shapes and leaves the mechanics to OpenAI.

What does one reuse of a 100K-token prefix cost?

Take a video-agent thread whose stable prefix (system instructions, tool list, brand brief and a long shot list) is 100K tokens. Multiply the changelog rates by 0.1 (100K is a tenth of a million tokens) and the table below falls out. These are list-price arithmetic, not measured bills.

The write premium is $0.05 on this prefix, and each read saves $0.19 against fresh input. One reuse pays the premium back almost four times over. The case where caching loses is the one-shot prefix: a single turn that writes and never reads costs $0.25 instead of $0.20.

Cost of a prefix is only part of a turn, since output is billed at $10 per 1M tokens on the same page. A turn that writes a 1,500-token shot list costs $0.015 in output, which is small next to a 100K-token prefix sent fresh at $0.20. That asymmetry is why prefix handling, not answer length, decides the bill on long threads.

GPT-6.1 Sol list rates applied to a 100K-token prefix (read 2026-10-03)
PatternCalculationPrefix cost
Send fresh, 1 turn100K x $2.00 per 1M$0.20
Send fresh, 2 turns2 x $0.20$0.40
Write once, read once$0.25 write + $0.01 read$0.26
Write once, read 9 times$0.25 + 9 x $0.01$0.34
Send fresh, 10 turns10 x $0.20$2.00

Which video-agent patterns reuse a prefix?

Reuse needs the early part of the prompt to stay byte-identical while later turns append. Three common shapes fit that.

  • A long thread where the brief and tool list stay fixed and each turn adds a render result: every turn after the first reads the prefix.
  • A batch of hook variants that share one brand brief but differ only in the last instruction: the first variant writes, the rest read.
  • A review pass that re-reads the same shot list with a different question each time.

Which patterns do not benefit?

Anything that edits the front of the prompt breaks the match: rewriting the system text each turn, shuffling the tool list, or inserting a timestamp near the top. A one-off job (one script in, one render out) writes a cache entry nobody reads, which at $2.50 is the only case that costs more than skipping the cache. The practical rule is to put the stable material first, the volatile material last, and expect savings only from the second turn onward.

How do I see what a Sume run actually spent?

Sume's Formats API lets you name the orchestrating model with the model field on Calling a Format; the id must be in the Agents catalog, and the receipt echoes the id that ran. gpt-6.1-sol is a catalog id in the Sume codebase. The run receipt carries usage.debited_usd_micros, which includes the agent's own LLM turn, so the receipt is where to compare a first turn against a continued one.

What Sume does not document is how it prices a vendor cache write or a cached read inside that debit. I found nothing on cache-write handling in the docs, so treat the table above as the vendor's list arithmetic and read your receipts for the number Sume charges. A practical test: run the same Format twice with previous_run_id and compare the two debited_usd_micros values.

curl -X POST https://api.sume.com/v1/formats/sume/sume-video-hook/runs \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: cache-test-001" \
  -d '{"instruction":"Three hook ideas for a vitamin C serum",
       "model":"gpt-6.1-sol",
       "generation_spend_cap_usd":5}'

Sources

Related posts

More in Models

All Models posts

Written by Sume