Cloudflare AI Search bills from Nov 1: split retrieval from renders

Cloudflare's October 1 changelog makes AI Search GA with usage billing from November 1, 2026. How to keep retrieval costs separate from Sume render costs.

4 min readSume
All posts

Cloudflare's changelog says AI Search became generally available on October 1, 2026 and that usage billing starts on November 1, 2026. If an agent of yours looks up brand or product facts through AI Search before it generates media, a new retrieval line appears on your Cloudflare invoice next to the generation cost you already track.

This post is about bookkeeping, not about AI Search itself. The entry I read gives the dates and does not state a price, so I do not quote one.

What the October 1 and 2 entries say

All items are from Cloudflare's developer changelog, read on 2026-10-03.

Cloudflare changelog items, October 1 and 2, 2026 (read 2026-10-03)
DateItemDetail in the entry
Oct 1AI SearchGenerally available; usage billing starts Nov 1, 2026
Oct 1ArtifactsOpen beta
Oct 1Basin (formerly Data Platform)Generally available
Oct 1Workers OAuth ProviderVersion 1.0
Oct 2Web Search APIBeta, for agents

A calendar item, not a migration

Nothing changes for a Sume integration on November 1. Sume jobs are billed by Sume, and the Sume API reference lists GET /v1/balance for the USD available balance and GET /v1/usage for ledger entries such as reservations, captures, refunds and top-ups. The new Cloudflare charge sits outside both.

What can go wrong is attribution. If one agent run does a retrieval, then three image jobs, then a video job, and you only see a combined spend graph, you cannot tell whether a cost rise came from more lookups or from more renders. Keep the two ledgers separate and join them by a run identifier you control.

Join the two on your own run id

Sume jobs accept metadata that is stored on the job and is not sent to the provider, per the video docs. Put your own run id there, and write the same id beside each retrieval record you keep. A weekly report can then group both sides by run.

from collections import defaultdict


def by_run(retrieval_rows, render_rows):
    """Each row: {"run": str, "usd": float}."""
    totals = defaultdict(lambda: {"retrieval": 0.0, "render": 0.0})
    for row in retrieval_rows:
        totals[row["run"]]["retrieval"] += row["usd"]
    for row in render_rows:
        totals[row["run"]]["render"] += row["usd"]
    return dict(totals)


if __name__ == "__main__":
    print(by_run([{"run": "r1", "usd": 0.01}], [{"run": "r1", "usd": 0.4}, {"run": "r2", "usd": 0.2}]))

Checklist before November 1

  • Find every agent that calls AI Search and note how often it does per run.
  • Read the live Cloudflare pricing page for the unit price; do not rely on a blog post.
  • Add a run id to Sume job metadata so renders can be grouped.
  • Decide a monthly cap for retrieval separate from the render cap.
  • Check the 402 insufficient_credits path in your agent so a drained Sume balance does not look like a retrieval failure.

Why not wait for the invoice

A first invoice is a poor place to learn how many lookups an agent makes. Add a counter now, while the service is not yet billed, and you will have a month of baseline data by November 1. If an agent loops on retrieval until it is satisfied, the counter will show it before it costs anything.

The same counter helps the other way: if retrieval turns out to be a tiny share, you can stop worrying and keep your attention on render spend, which the Sume ledger already breaks into reservations, captures and refunds.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume