Cached input price for agent LLMs: GPT-6.1 Sol to DeepSeek
Cached input per million tokens for GPT-6.1 Sol, Opus 5.5, Sonnet 5.5, Grok 4.7 and DeepSeek Flash, and why it sets the cost of a long video agent thread.

On the vendors' own pages, a cached input token costs $0.10 per million on GPT-6.1 Sol, $0.20 on Claude Opus 5.5, $0.20 on Claude Sonnet 5.5, $0.50 on Grok 4.7 below 200k tokens, and between $0.003 and $0.006 on DeepSeek's deepseek-flash. A video agent resends its whole thread on every turn, so that one number decides what a long project costs more than the output price does.
Prices below come from GPT-6.1 Sol, GPT-6 Sol, Anthropic's models overview and pricing page, Grok 4.7 and DeepSeek's Models & Pricing, all read on 2026-10-02. They are vendor list prices, not what Sume charges.
What does a cached input token cost on each model?
Cached input is the price for a prompt prefix the vendor has already processed and stores. The table puts it next to the normal input price, so you can see the discount as a share of list.
- DeepSeek lists a range because off-peak rates are lower; its page puts peak hours at 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.
- Anthropic prices a cache hit at 0.1x base input, except 0.05x on Opus 5.5. A five-minute cache write costs 1.25x base input, and a one-hour write costs 2x.
- OpenAI's GPT-6.1 Sol page also lists cache writes at $2.50 per million; the GPT-6 Sol page does not list a write price.
| Model | Input | Cached input | Cached as share of input | Output |
|---|---|---|---|---|
| GPT-6.1 Sol | $2 | $0.10 | 5% | $10 |
| GPT-6 Sol | $2 | $0.20 | 10% | $10 |
| Claude Opus 5.5 | $4 | $0.20 | 5% | $20 |
| Claude Sonnet 5.5 | $2 | $0.20 | 10% | $10 |
| Grok 4.7 (prompt under 200k) | $2 | $0.50 | 25% | $6 |
| DeepSeek deepseek-flash | $0.15 to $0.30 | $0.003 to $0.006 | 2% | $0.60 to $1.20 |
Why does it matter more for agents than for chat?
A chat turn sends a short prompt. An agent turn sends the system prompt, the tool definitions, every earlier tool result, and the plan so far. In a video job that includes shot lists, job ids, and notes on each clip. Each new turn repeats the same prefix and adds a little at the end.
Take a 150,000-token prefix that is re-read on 20 turns, which is 3 million cached input tokens. At the cached rates above that costs $0.30 on GPT-6.1 Sol, $0.60 on Opus 5.5 or Sonnet 5.5, $1.50 on Grok 4.7, and roughly one to two cents on DeepSeek. The same 3 million tokens at the uncached input price would be $6 on GPT-6.1 Sol, $12 on Opus 5.5, $6 on Sonnet 5.5 and Grok 4.7, and $0.45 to $0.90 on DeepSeek.
This is arithmetic on list prices, not a measured result. It assumes every re-read hits the cache, which depends on the prefix staying byte-identical and on the cache lifetime. Anthropic's five-minute default expires if your agent waits on a long render between turns.
Is a lower cached price always the cheaper agent?
No. The cached rate only prices the part of the prompt that repeats. New tokens pay full input, and every output token pays the output rate, which includes thinking tokens on reasoning models. A model with a cheap cache and a slow, wordy plan can cost more per finished video than a pricier one that plans in fewer turns.
Grok 4.7 has one more wrinkle. Its page says a request whose prompt reaches 200k tokens is billed at the higher rate for all tokens: $4 input, $1 cached, $12 output. The table shows only the lower tier, and the 200k jump gets its own post.
What does Sume show you about this?
Sume does not expose a cache setting. You pick an orchestrator with the model field on a Format run, and the docs do not describe a cache setting or promise a cache hit rate, so do not budget on one.
What you can read is the result. The run receipt has a usage object. billable_amount_usd_micros is the generation spend counted against your cap and excludes the agent's own LLM turn. debited_usd_micros is what the wallet deducted, including the turn's LLM row. Compare that figure across the same job on two models, and you have the real cached-versus-uncached answer for your workload.
In the repo, GPT-6.1 Sol has its own rate card instead of reusing GPT-6 Sol's, because its cached input is half of GPT-6 Sol's. Whatever the list price on a vendor page, Sume meters at the rates on its API pricing page, so check that page before you copy any number here into a quote.
How do you pick one for a video agent?
Start from the thread, not the model. If your jobs are short, with a few turns and a small prefix, the cached price barely registers and you should choose on quality and latency. If you continue the same conversation across many scenes with previous_run_id, the cache matters and the 5% models, GPT-6.1 Sol and Opus 5.5, have an edge on paper. Run the same input on two models and read debited_usd_micros before you commit.
Sources
- OpenAI API: GPT-6.1 Sol model page (read 2026-10-02)
- OpenAI API: GPT-6 Sol model page (read 2026-10-02)
- Anthropic: Models overview (read 2026-10-02)
- Anthropic: Pricing (read 2026-10-02)
- xAI Docs: Grok 4.7 (read 2026-10-02)
- DeepSeek API Docs: Models & Pricing (read 2026-10-02)
- Create a run (Formats API)
- Runs and results (Formats API)
Related posts
More in Models
- Context window and max output for agent LLMs, compared
Context and output limits for GPT-6.1 Sol, Opus 5.5, Sonnet 5.5, Grok 4.7 and DeepSeek Flash from vendor pages, and what they mean for a long video thread.
- Dialogue in Korean, Japanese or Spanish: Kling 3.0 vs Seedance 2.5
Which spoken languages Kling 3.0 Omni and Seedance 2.5 list, what Sume's generate_audio flag does, and how to test a Korean line before a full render.
- Which AI video model for a 15-second single take on Sume?
Kling 3.0, Wan 3.0, MiniMax H3 and Seedance 2 reach 15 seconds on Sume; Auto and Grok stop at 10. Seedance 2.5 and Wan 3.0 go to 30. Table of limits.
- POST /v1/avatar-1.0/fabric is test-only: use veed/fabric-1.0 instead
The experimental avatar-1.0/fabric route is a temporary comparison endpoint that may be removed. How it differs from image-to-video and what to build on.
Written by Sume