Clef context: 65,536 on Cloudflare, 16,384 on the model card?

Cloudflare's Workers AI page lists 65,536 tokens of context for Clef; its Hugging Face card says 16,384 by default. Budget from the smaller. Cost math inside.

5 min readSume
All posts

Two Cloudflare pages give two numbers for Clef. The Workers AI docs list a context window of 65,536 tokens. The Hugging Face model card says the model accepts up to 16,384 tokens by default, adjustable with the max_length parameter. If you do not control that parameter, budget your decision prompts against 16,384.

What each vendor page says

Clef is Cloudflare's open decision model. The Workers AI page, read on 2026-10-08, describes a 27B multimodal model that turns a state and a schema of typed questions into decisions. It takes text, JSON, images, or video, and returns a probability for every option. The card on Hugging Face adds the Apache-2.0 license and says the model outputs one logit per allowed option per question, with a softmax per question to get probabilities.

Clef context and price, from two Cloudflare pages (read 2026-10-08)
SourceContext figureOther facts
Workers AI docs65,536 tokens$0.24 per million input tokens; ids @cf/cloudflare/clef and @cf/cloudflare/clef-flash
Hugging Face card16,384 tokens by default; adjustable via max_length27B, Apache-2.0

Reading the difference

The pages are not in direct conflict. One names the model's window, and the other names a default for a loading parameter. I cannot tell from the pages which one governs a Workers AI call, so plan for the smaller. If a decision needs more state than that, summarize the state first.

What a decision prompt costs

Using only the $0.24 per million input rate from the Workers AI page, a full 16,384-token prompt costs 16,384 x 0.24 / 1,000,000 = about $0.0039. A thousand of them cost about $3.93. At the 65,536 figure a thousand full prompts would cost about $15.73. Real prompts are usually far shorter than a full window, so these are ceilings.

For comparison, a Sume Agent Completion needs a generation_spend_cap_usd on every call, and the docs say the cap is the maximum generation spend on that run. A gate that costs a few cents per thousand checks is cheap next to a run you would cancel.

What to put in a state block for a video agent

The points that matter here, in the order you will hit them:

  • The run status, the last error code, and the count of artifacts so far.
  • The requested format and duration, as short strings.
  • The remaining cap in dollars, so the decision can see the limit.
  • No signed URLs and no API keys; the Sume safe-automation guidance lists both as unsafe to log.

Sume's side

Sume's agent model catalog in the repository, read on 2026-10-08, has no Clef row, and the Agent Completions model field accepts only sume-agent. A Clef gate lives in your own service, before you call the Sume API.

When the state is too big

If a run's state serializes to 40,000 tokens, it does not fit in one 16,384-token prompt. ceil(40,000 / 16,384) = 3 prompts would be needed, and splitting a decision across prompts loses the cross-references the question may depend on. Summarize first. Keep identifiers, statuses, and counts, and drop logs. A decision model answers typed questions about a state block, so the state should be small and structured.

Sources

Related posts

More in Models

All Models posts

Written by Sume