Prompt cache diagnostics are GA: why did my video agent miss?

Anthropic took cache diagnostics out of beta on Sep 23 and OpenAI made them GA on Sep 8. What each reports, and where Sume's run receipt fits in.

5 min readSume
All posts

When a long agent thread suddenly gets slower and dearer, the likely cause is a prompt-cache miss, and as of this month both Anthropic and OpenAI have a supported way to ask why. Anthropic's cache diagnostics left beta on September 23, and OpenAI made Prompt Cache Diagnostics generally available in the Responses API on September 8 for GPT-5.6 and later models.

Both statements come from the vendors' own pages: the Claude Platform release notes and the OpenAI API changelog, read on 2026-10-03. Neither page lists the individual miss reasons, so this post covers how to switch the feature on and what to do with the answer, not a catalogue of causes.

What changed on the Anthropic side?

The release notes give four dated entries. September 9: fingerprints are stored only when the request includes a diagnostics object; a request carrying only the beta header stores none, and previous_message_id reports previous_message_not_found if the earlier message is missing. September 18 and 23: the diagnostics field is always present on POST /v1/messages responses, null when you did not ask. September 23: the feature no longer needs the cache-diagnosis-2026-04-07 beta header; you opt in by including the diagnostics object on the request.

What changed on the OpenAI side?

One line in the changelog: Prompt Cache Diagnostics is generally available in the Responses API for GPT-5.6+ models, to compare cache reuse and identify miss reasons. The entry sits beside the GPT-6 and GPT-6.1 launches, which the same page prices with cached-input rates ($0.20 per 1M for GPT-6 Sol, $0.10 for GPT-6.1 Sol), so a miss is a visible cost line.

Cache diagnostics as the vendors describe them (read 2026-10-03)
VendorWhereOpt inDate
AnthropicMessages APIInclude a diagnostics object on the requestOut of beta Sep 23, 2026
OpenAIResponses API, GPT-5.6+ modelsGenerally available (details on OpenAI's page)GA Sep 8, 2026

What do I do with a miss in a video-agent thread?

Cache problems in agent threads are usually self-inflicted: something near the top of the prompt changed. Typical culprits are a tool list that is reordered, a timestamp or run id placed in the system text, or a brand brief that is regenerated per turn. Use diagnostics on a small replay (two identical requests, then one with a single edit) so you learn which change flips the result before you touch production prompts.

Two habits make the readout meaningful. Keep the stable instructions first and the per-turn material last, and keep the model id fixed within a thread, since a different model is a different cache. The second habit is one you can log: record the model id next to every request.

Where does Sume fit?

Sume does not expose either vendor's diagnostics object; I found no cache-diagnostics field in the Sume docs. What it does give you is a per-run receipt. Calling a Format lets you set model to an Agents catalog id and says the receipt echoes the id that ran, and Format runs documents usage.debited_usd_micros, which includes the agent's own LLM turn. A continued run (previous_run_id) is the same conversation, so two receipts from one thread are a coarse before-and-after for cost.

If you need the vendor's miss reason itself, call that vendor's API directly with the diagnostics option on. If you only need to know whether your thread got dearer, compare receipts.

curl https://api.sume.com/v1/format-runs/$RUN_ID \
  -H "Authorization: Bearer $SUME_API_KEY" \
  | jq '{model, debited: .usage.debited_usd_micros, final: .usage.final}'

What is a sensible replay test?

Build a three-request experiment. Request one is your real prompt. Request two is byte-identical to it. Request three changes exactly one early element, such as reordering two tools. With diagnostics on, request two should show reuse and request three should show a miss with a reason attached; if request two also misses, something you thought was static (a timestamp, a random id, whitespace) is not.

Run the experiment per model, since a cache belongs to a model. When you find the unstable element, fix it at the source: generate ids and dates after the stable prefix, sort tool lists deterministically, and keep brand briefs as stored text rather than regenerated text. Repeat the replay after any prompt edit that touches the top of the thread.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume