Clef-omni reads video: a cheap check before a second video job
Cloudflare's Clef-omni takes audio and video input at $0.15 per million tokens. Sume does not offer Clef, so this is a pre-check pattern to run outside Sume.

Clef-omni, announced by Cloudflare on October 9, 2026, is an open-weight decision model that reads text, images, audio and video in one request and returns a probability for each allowed answer. Cloudflare prices it at $0.15 per million input tokens and says a 21-second video with sound takes about 1.5 seconds. Sume does not list Clef, so use it outside Sume as a cheap check before you pay for another generation.
What does Clef-omni do?
Cloudflare's post says Clef-omni accepts wav and mp3 audio, mp4 and webm video, images and text in a single API call. It runs a prefill pass over the whole payload and scores every valid option for each question at once, so the output is a set of confidence values, not free text. It is built on Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model, with the text-to-speech parts removed, and the weights are on Hugging Face.
How fast and how cheap are the three tiers?
The post gives these numbers for the launch. Latency figures are medians Cloudflare reports for short payloads.
| Tier | Input price per million tokens | Notes from the post |
|---|---|---|
| Clef-omni | $0.15 | Audio, video, image and text; about 1.5 s for a 21 s clip with sound |
| Clef | $0.24 | 64k context; 1.7x median speedup at about 800 tokens |
| Clef-flash | $0.038 (was $0.09) | Hosted context cut from 64k to 24k |
Where would a video check fit with Sume?
A Sume video job is asynchronous and billed per output second, so a rejected clip costs a full render. A fast yes-or-no check on an input photo or a finished clip, run in your own code before you submit the next job, can stop obvious misses: the wrong product in frame, text where there should be none, a person you did not intend. Sume does not run Clef for you: the Sume repo has no Clef model, so you would call Cloudflare yourself.
- Ask closed questions with a fixed answer list, such as product visible: yes or no.
- Only submit the next Sume video job when the confidence is above a threshold you choose.
- Keep the threshold in your code; a confidence score is not a spend cap.
What should you verify first?
Confirm context limits and pricing on Cloudflare's model page before you build, because the hosted Clef-flash context differs from the self-hosted weights. Then test the gate on a few dozen of your own clips: a decision model can be confidently wrong on your particular catalog, and the cheapest check is the one you have measured.
Sources
Related posts
More in Models
- FastH3 Trim: MiniMax H3 on an 8 GB GPU, and who may use it
FastH3 Trim prunes MiniMax H3 to 42 blocks and 8 steps. File sizes, the 8 GB claim, the quality trade, and what the H3 Community License allows.
- Image releases of Oct 6-10, 2026: which have a Sume model id
Nano Banana 2.1 is google/nano-banana-2.1 on Sume. Qwen-Image-2.1-Turbo and Kroma have no id; qwen/qwen-image is the nearest Qwen row. Check yours in code.
- Kandinsky 6.0 video with sound: is it on Sume, and what to use
Kandinsky 6.0 makes video and audio together, MIT-licensed. Sume does not list it. Here is what Sume lists for synced sound, and when to self-host Kandinsky.
- Kandinsky 6.0 Video is MIT: can you use the clips commercially?
Kandinsky 6.0 Video Pro (29B) and Lite (3B) are MIT-licensed with joint audio. What MIT covers, what to check, and the hosted alternative on Sume.
Written by Sume