Prompt cache diagnostics are GA: why did my video agent miss?
Anthropic took cache diagnostics out of beta on Sep 23 and OpenAI made them GA on Sep 8. What each reports, and where Sume's run receipt fits in.

When a long agent thread suddenly gets slower and dearer, the likely cause is a prompt-cache miss, and as of this month both Anthropic and OpenAI have a supported way to ask why. Anthropic's cache diagnostics left beta on September 23, and OpenAI made Prompt Cache Diagnostics generally available in the Responses API on September 8 for GPT-5.6 and later models.
Both statements come from the vendors' own pages: the Claude Platform release notes and the OpenAI API changelog, read on 2026-10-03. Neither page lists the individual miss reasons, so this post covers how to switch the feature on and what to do with the answer, not a catalogue of causes.
What changed on the Anthropic side?
The release notes give four dated entries. September 9: fingerprints are stored only when the request includes a diagnostics object; a request carrying only the beta header stores none, and previous_message_id reports previous_message_not_found if the earlier message is missing. September 18 and 23: the diagnostics field is always present on POST /v1/messages responses, null when you did not ask. September 23: the feature no longer needs the cache-diagnosis-2026-04-07 beta header; you opt in by including the diagnostics object on the request.
What changed on the OpenAI side?
One line in the changelog: Prompt Cache Diagnostics is generally available in the Responses API for GPT-5.6+ models, to compare cache reuse and identify miss reasons. The entry sits beside the GPT-6 and GPT-6.1 launches, which the same page prices with cached-input rates ($0.20 per 1M for GPT-6 Sol, $0.10 for GPT-6.1 Sol), so a miss is a visible cost line.
| Vendor | Where | Opt in | Date |
|---|---|---|---|
| Anthropic | Messages API | Include a diagnostics object on the request | Out of beta Sep 23, 2026 |
| OpenAI | Responses API, GPT-5.6+ models | Generally available (details on OpenAI's page) | GA Sep 8, 2026 |
What do I do with a miss in a video-agent thread?
Cache problems in agent threads are usually self-inflicted: something near the top of the prompt changed. Typical culprits are a tool list that is reordered, a timestamp or run id placed in the system text, or a brand brief that is regenerated per turn. Use diagnostics on a small replay (two identical requests, then one with a single edit) so you learn which change flips the result before you touch production prompts.
Two habits make the readout meaningful. Keep the stable instructions first and the per-turn material last, and keep the model id fixed within a thread, since a different model is a different cache. The second habit is one you can log: record the model id next to every request.
Where does Sume fit?
Sume does not expose either vendor's diagnostics object; I found no cache-diagnostics field in the Sume docs. What it does give you is a per-run receipt. Calling a Format lets you set model to an Agents catalog id and says the receipt echoes the id that ran, and Format runs documents usage.debited_usd_micros, which includes the agent's own LLM turn. A continued run (previous_run_id) is the same conversation, so two receipts from one thread are a coarse before-and-after for cost.
If you need the vendor's miss reason itself, call that vendor's API directly with the diagnostics option on. If you only need to know whether your thread got dearer, compare receipts.
curl https://api.sume.com/v1/format-runs/$RUN_ID \
-H "Authorization: Bearer $SUME_API_KEY" \
| jq '{model, debited: .usage.debited_usd_micros, final: .usage.final}'What is a sensible replay test?
Build a three-request experiment. Request one is your real prompt. Request two is byte-identical to it. Request three changes exactly one early element, such as reordering two tools. With diagnostics on, request two should show reuse and request three should show a miss with a reason attached; if request two also misses, something you thought was static (a timestamp, a random id, whitespace) is not.
Run the experiment per model, since a cache belongs to a model. When you find the unstable element, fix it at the source: generate ids and dates after the stable prefix, sort tool lists deterministically, and keep brand briefs as stored text rather than regenerated text. Repeat the replay after any prompt edit that touches the top of the thread.
Sources
Related posts
More in Developers
- Pydantic model to Sume output_schema: extra forbid, no defaults
Turn a Pydantic v2 model into a valid Sume Format output_schema: extra=forbid, nullable instead of defaults, and the SumeMediaFile reference. Tested.
- Unit test Sume webhook signature checks in pytest (Python)
A pytest file for HMAC webhook verification: tamper, rotation header, stale timestamp and empty secret, written against Sume's sume-v1 scheme.
- Log x-sume-request-id and Idempotency-Key on every call (Python)
A requests response hook that writes one JSON log line per Sume call: x-sume-request-id, idempotency key, error code and rate-limit headers. Tested.
- Does a queue_full 429 charge me? Sume's reservation rules
A Sume 429 queue_full means the workspace has no accepted-job capacity left. The failed admission releases its reservation; retry with the same Idempotency-Key.
Written by Sume