Cloudflare AI Search bills from Nov 1: split retrieval from renders
Cloudflare's October 1 changelog makes AI Search GA with usage billing from November 1, 2026. How to keep retrieval costs separate from Sume render costs.

Cloudflare's changelog says AI Search became generally available on October 1, 2026 and that usage billing starts on November 1, 2026. If an agent of yours looks up brand or product facts through AI Search before it generates media, a new retrieval line appears on your Cloudflare invoice next to the generation cost you already track.
This post is about bookkeeping, not about AI Search itself. The entry I read gives the dates and does not state a price, so I do not quote one.
What the October 1 and 2 entries say
All items are from Cloudflare's developer changelog, read on 2026-10-03.
| Date | Item | Detail in the entry |
|---|---|---|
| Oct 1 | AI Search | Generally available; usage billing starts Nov 1, 2026 |
| Oct 1 | Artifacts | Open beta |
| Oct 1 | Basin (formerly Data Platform) | Generally available |
| Oct 1 | Workers OAuth Provider | Version 1.0 |
| Oct 2 | Web Search API | Beta, for agents |
A calendar item, not a migration
Nothing changes for a Sume integration on November 1. Sume jobs are billed by Sume, and the Sume API reference lists GET /v1/balance for the USD available balance and GET /v1/usage for ledger entries such as reservations, captures, refunds and top-ups. The new Cloudflare charge sits outside both.
What can go wrong is attribution. If one agent run does a retrieval, then three image jobs, then a video job, and you only see a combined spend graph, you cannot tell whether a cost rise came from more lookups or from more renders. Keep the two ledgers separate and join them by a run identifier you control.
Join the two on your own run id
Sume jobs accept metadata that is stored on the job and is not sent to the provider, per the video docs. Put your own run id there, and write the same id beside each retrieval record you keep. A weekly report can then group both sides by run.
from collections import defaultdict
def by_run(retrieval_rows, render_rows):
"""Each row: {"run": str, "usd": float}."""
totals = defaultdict(lambda: {"retrieval": 0.0, "render": 0.0})
for row in retrieval_rows:
totals[row["run"]]["retrieval"] += row["usd"]
for row in render_rows:
totals[row["run"]]["render"] += row["usd"]
return dict(totals)
if __name__ == "__main__":
print(by_run([{"run": "r1", "usd": 0.01}], [{"run": "r1", "usd": 0.4}, {"run": "r2", "usd": 0.2}]))Checklist before November 1
- Find every agent that calls AI Search and note how often it does per run.
- Read the live Cloudflare pricing page for the unit price; do not rely on a blog post.
- Add a run id to Sume job
metadataso renders can be grouped. - Decide a monthly cap for retrieval separate from the render cap.
- Check the
402 insufficient_creditspath in your agent so a drained Sume balance does not look like a retrieval failure.
Why not wait for the invoice
A first invoice is a poor place to learn how many lookups an agent makes. Add a counter now, while the service is not yet billed, and you will have a month of baseline data by November 1. If an agent loops on retrieval until it is satisfied, the counter will show it before it costs anything.
The same counter helps the other way: if retrieval turns out to be a tiny share, you can stop worrying and keep your attention on render spend, which the Sume ledger already breaks into reservations, captures and refunds.
Sources
Related posts
More in Developers
- Codex 0.160 agent history 'Show more': recover Sume jobs by job list
Codex CLI 0.160.0 adds Show more pagination to agent command center history. If a thread scrolls away, Sume's GET /v1/jobs list still holds every job.
- Cursor Security Review bot on a Sume webhook handler: what to find
Cursor added a Security Review bot on Sep 23. A webhook handler for Sume should pass seven checks: raw body, timestamp window, rotation, empty secret and more.
- Cursor self-hosted machines can't take Sume webhooks on a private URL
Sume rejects localhost, private-network and non-HTTPS webhook URLs. An agent on a self-hosted machine should poll status_url or use a public HTTPS receiver.
- Demand Gen copy limits: 40-character headlines, one at 30 or fewer
Demand Gen allows 40-character headlines (one must be 30 or fewer), 90-character descriptions and 10-60 second videos. A checker script plus the Sume lengths.
Written by Sume