Cached input price for agent LLMs: GPT-6.1 Sol to DeepSeek

Cached input per million tokens for GPT-6.1 Sol, Opus 5.5, Sonnet 5.5, Grok 4.7 and DeepSeek Flash, and why it sets the cost of a long video agent thread.

5 min readSume
All posts

On the vendors' own pages, a cached input token costs $0.10 per million on GPT-6.1 Sol, $0.20 on Claude Opus 5.5, $0.20 on Claude Sonnet 5.5, $0.50 on Grok 4.7 below 200k tokens, and between $0.003 and $0.006 on DeepSeek's deepseek-flash. A video agent resends its whole thread on every turn, so that one number decides what a long project costs more than the output price does.

Prices below come from GPT-6.1 Sol, GPT-6 Sol, Anthropic's models overview and pricing page, Grok 4.7 and DeepSeek's Models & Pricing, all read on 2026-10-02. They are vendor list prices, not what Sume charges.

What does a cached input token cost on each model?

Cached input is the price for a prompt prefix the vendor has already processed and stores. The table puts it next to the normal input price, so you can see the discount as a share of list.

  • DeepSeek lists a range because off-peak rates are lower; its page puts peak hours at 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.
  • Anthropic prices a cache hit at 0.1x base input, except 0.05x on Opus 5.5. A five-minute cache write costs 1.25x base input, and a one-hour write costs 2x.
  • OpenAI's GPT-6.1 Sol page also lists cache writes at $2.50 per million; the GPT-6 Sol page does not list a write price.
Input, cached input and output per million tokens, vendor pages, read 2026-10-02
ModelInputCached inputCached as share of inputOutput
GPT-6.1 Sol$2$0.105%$10
GPT-6 Sol$2$0.2010%$10
Claude Opus 5.5$4$0.205%$20
Claude Sonnet 5.5$2$0.2010%$10
Grok 4.7 (prompt under 200k)$2$0.5025%$6
DeepSeek deepseek-flash$0.15 to $0.30$0.003 to $0.0062%$0.60 to $1.20

Why does it matter more for agents than for chat?

A chat turn sends a short prompt. An agent turn sends the system prompt, the tool definitions, every earlier tool result, and the plan so far. In a video job that includes shot lists, job ids, and notes on each clip. Each new turn repeats the same prefix and adds a little at the end.

Take a 150,000-token prefix that is re-read on 20 turns, which is 3 million cached input tokens. At the cached rates above that costs $0.30 on GPT-6.1 Sol, $0.60 on Opus 5.5 or Sonnet 5.5, $1.50 on Grok 4.7, and roughly one to two cents on DeepSeek. The same 3 million tokens at the uncached input price would be $6 on GPT-6.1 Sol, $12 on Opus 5.5, $6 on Sonnet 5.5 and Grok 4.7, and $0.45 to $0.90 on DeepSeek.

This is arithmetic on list prices, not a measured result. It assumes every re-read hits the cache, which depends on the prefix staying byte-identical and on the cache lifetime. Anthropic's five-minute default expires if your agent waits on a long render between turns.

Is a lower cached price always the cheaper agent?

No. The cached rate only prices the part of the prompt that repeats. New tokens pay full input, and every output token pays the output rate, which includes thinking tokens on reasoning models. A model with a cheap cache and a slow, wordy plan can cost more per finished video than a pricier one that plans in fewer turns.

Grok 4.7 has one more wrinkle. Its page says a request whose prompt reaches 200k tokens is billed at the higher rate for all tokens: $4 input, $1 cached, $12 output. The table shows only the lower tier, and the 200k jump gets its own post.

What does Sume show you about this?

Sume does not expose a cache setting. You pick an orchestrator with the model field on a Format run, and the docs do not describe a cache setting or promise a cache hit rate, so do not budget on one.

What you can read is the result. The run receipt has a usage object. billable_amount_usd_micros is the generation spend counted against your cap and excludes the agent's own LLM turn. debited_usd_micros is what the wallet deducted, including the turn's LLM row. Compare that figure across the same job on two models, and you have the real cached-versus-uncached answer for your workload.

In the repo, GPT-6.1 Sol has its own rate card instead of reusing GPT-6 Sol's, because its cached input is half of GPT-6 Sol's. Whatever the list price on a vendor page, Sume meters at the rates on its API pricing page, so check that page before you copy any number here into a quote.

How do you pick one for a video agent?

Start from the thread, not the model. If your jobs are short, with a few turns and a small prefix, the cached price barely registers and you should choose on quality and latency. If you continue the same conversation across many scenes with previous_run_id, the cache matters and the 5% models, GPT-6.1 Sol and Opus 5.5, have an edge on paper. Run the same input on two models and read debited_usd_micros before you commit.

Sources

Related posts

More in Models

All Models posts

Written by Sume