Clef context: 65,536 on Cloudflare, 16,384 on the model card?
Cloudflare's Workers AI page lists 65,536 tokens of context for Clef; its Hugging Face card says 16,384 by default. Budget from the smaller. Cost math inside.

Two Cloudflare pages give two numbers for Clef. The Workers AI docs list a context window of 65,536 tokens. The Hugging Face model card says the model accepts up to 16,384 tokens by default, adjustable with the max_length parameter. If you do not control that parameter, budget your decision prompts against 16,384.
What each vendor page says
Clef is Cloudflare's open decision model. The Workers AI page, read on 2026-10-08, describes a 27B multimodal model that turns a state and a schema of typed questions into decisions. It takes text, JSON, images, or video, and returns a probability for every option. The card on Hugging Face adds the Apache-2.0 license and says the model outputs one logit per allowed option per question, with a softmax per question to get probabilities.
| Source | Context figure | Other facts |
|---|---|---|
| Workers AI docs | 65,536 tokens | $0.24 per million input tokens; ids @cf/cloudflare/clef and @cf/cloudflare/clef-flash |
| Hugging Face card | 16,384 tokens by default; adjustable via max_length | 27B, Apache-2.0 |
Reading the difference
The pages are not in direct conflict. One names the model's window, and the other names a default for a loading parameter. I cannot tell from the pages which one governs a Workers AI call, so plan for the smaller. If a decision needs more state than that, summarize the state first.
What a decision prompt costs
Using only the $0.24 per million input rate from the Workers AI page, a full 16,384-token prompt costs 16,384 x 0.24 / 1,000,000 = about $0.0039. A thousand of them cost about $3.93. At the 65,536 figure a thousand full prompts would cost about $15.73. Real prompts are usually far shorter than a full window, so these are ceilings.
For comparison, a Sume Agent Completion needs a generation_spend_cap_usd on every call, and the docs say the cap is the maximum generation spend on that run. A gate that costs a few cents per thousand checks is cheap next to a run you would cancel.
What to put in a state block for a video agent
The points that matter here, in the order you will hit them:
- The run status, the last error code, and the count of artifacts so far.
- The requested format and duration, as short strings.
- The remaining cap in dollars, so the decision can see the limit.
- No signed URLs and no API keys; the Sume safe-automation guidance lists both as unsafe to log.
Sume's side
Sume's agent model catalog in the repository, read on 2026-10-08, has no Clef row, and the Agent Completions model field accepts only sume-agent. A Clef gate lives in your own service, before you call the Sume API.
When the state is too big
If a run's state serializes to 40,000 tokens, it does not fit in one 16,384-token prompt. ceil(40,000 / 16,384) = 3 prompts would be needed, and splitting a decision across prompts loses the cross-references the question may depend on. Summarize first. Keep identifiers, statuses, and counts, and drop logs. A decision model answers typed questions about a state block, so the state should be small and structured.
Sources
Related posts
More in Models
- Do you have to credit an open video model? Label, notice, license
H3 requires a visible 'MiniMax H3' label in commercial products. Hunyuan only encourages 'Powered by'. LTX and Wan ask for notices and a license copy. Per file.
- Does a vertical 9:16 Omni Flash clip cost more than 16:9?
No: on Sume a 9:16 and a 16:9 Gemini Omni Flash 1.1 clip cost the same at each resolution and length. The side-by-side numbers and how to request each ratio.
- Does Sume offer Haiku 5.5, GLM 5.3 Flash or Mistral Large 4?
Checked against Sume's model catalog on 2026-10-08: Haiku 5.5 and GLM 5.3 Flash have enabled rows; Mistral Large 4, Clef and Strands Decider 2B have none.
- Does SynthID survive captions and a 9:16 crop on Omni clips?
Google says SynthID is built to survive cropping, filters and compression. What that means for a captioned or reframed Omni clip, and what Sume claims.
Written by Sume