Primary sourceAIArticle··6 min read

Cloudflare launches Clef-omni: a decision model that sees, hears, and decides

A week after its first open-weight models, Cloudflare adds audio and video, slashes Clef-flash's price by more than half, and speeds up Clef.

Cloudflare launches Clef-omni: a decision model that sees, hears, and decides
Source : Cloudflare · Cloudflare Blog · 9 October 2026View original ↗

In brief

Cloudflare expands its Clef family of decision models with Clef-omni, which processes text, images, audio, and video in a single call. Clef-flash drops to $0.038 per million input tokens, undercutting competitor Jev, at the cost of a smaller context window, while Clef gets up to 2× faster. The goal: turn classification and detection into a simple API building block for agentic workflows.

🍺 Bar-stool version

Cloudflare builds models that don't chit-chat: you ask them a closed question, they answer yes or no with a confidence score, full stop. Now they've given them ears and eyes, so you can send the photo, sound, and video of an air conditioner and ask if it's making a suspicious noise, without routing the job through three different models passing the baton. They also claim they decided on the project on a Friday night and trained it over the weekend, which says a lot about their weekends. What really matters: sorting, detecting, and moderating becomes as mundane as calling an API, and at this price, nobody has an excuse to let spam slip through anymore.

Key takeaways

  1. 1

    Clef-omni accepts audio (wav, mp3), video (mp4, webm), images, and text in a single API call, with no transcription pipeline or track-splitting required.

  2. 2

    The model is built on Qwen3-Omni-30B-A3B-Instruct (mixture-of-experts), of which Cloudflare keeps the comprehension backbone while stripping out speech synthesis.

  3. 3

    Clef is not an LLM: it doesn't generate tokens, it runs a prefill over the entire payload and scores each valid option via two-stage attention routing.

  4. 4

    Reported median latencies: about 130 ms for text, 150 ms for image, a few hundred milliseconds for audio, and 1.5 s for a 21-second video with sound.

  5. 5

    Clef-flash drops from $0.09 to $0.038 per million input tokens, but its hosted version sees its context cut from 64k to 24k tokens; Clef stays at $0.24 and Clef-omni comes in at $0.15.

  6. 6

    Clef gains 1.7× to 2× in median latency thanks to the switch to SGLang, with upstream integration planned for SGLang 0.5.22.

  7. 7

    Benchmarks show Clef-omni lagging behind Clef on several tasks, notably Home appliances (69.3 vs 82.95) and across all TypeSafe evaluations.

A multimodal decision model

Clef-omni is the third piece of the Clef family, launched a week earlier alongside Clef and Clef-flash. While Clef already accepted images and video frame grids, Clef-omni directly ingests audio and video files, on top of text and images.

Cloudflare frames this as a paradigm shift for a model category that has remained largely text-based since Jev, by TypeSafe, arrived. The example given in the documentation is telling: a photo, an audio recording, and a video of an installed appliance, with three yes/no questions about the visible label, the operating noise, and the fan rotation.

The stated benefit is simplification. No more chaining a speech-to-text model, a vision model, and a classifier: a single call is enough to decide based on any modality. The weights are published on Hugging Face.

Under the hood: no generation, just scoring

Clef-omni is built on Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model already capable of processing text, image, audio, and video within the same pipeline. Cloudflare keeps the comprehension backbone and discards the text-to-speech output components.

The model generates no tokens at all. It performs a fast prefill over the entire payload, with audio and video synchronized to frames, then extracts candidate values from its internal embeddings. Each valid option first retrieves its cues from the input, before field vectors perform cross-attention over the full context to compute confidence scores.

Training follows the same recipe as Clef: frozen Qwen3 backbone, LoRA adapters, cross-entropy loss with label smoothing combined with Brier score calibration. A built-in lexical grammar preserves the semantics of the options to guarantee schema-constrained outputs.

Mixed benchmark results

Clef-omni holds up well on several tests: 97.7 macro-F1 on CLINC150+OOS, 94.8 on BANKING77, 98.2 on BFCL, 57.8 on Amazon ESCI, matching or beating Clef and clearly ahead of Jev on some.

But multimodality comes at a cost. On Home appliances, Clef-omni drops to 69.3 versus 82.95 for Clef and 97.73 for Clef-flash. On When2Call, it scores 63.3 while Jev reaches 80.97. On PhishNChips, it falls to 73.2 versus 79.6 for Clef.

On TypeSafe's workflow evaluations (invoices, customer service, security incidents, agent trace observability), Clef-omni consistently underperforms Clef, and even falls below Jev in customer service (71.6 vs 76.0) and observability (65.8 vs 71.6). The omni model is thus a tool for multimodal use cases, not a universal replacement.

Clef-flash cheaper, but with less context

Clef-flash drops from $0.09 to $0.038 per million input tokens, making it cheaper than Jev according to Cloudflare. Clef stays at $0.24, Clef-omni starts at $0.15. The conversion of images and audio into tokens for billing purposes is detailed in the documentation.

The trade-off: the hosted version of Clef-flash sees its context window shrink from 64k to 24k tokens. Cloudflare justifies this based on usage data, with only 0.24% of requests exceeding 24k tokens.

The weights on Hugging Face remain unchanged and support 256k tokens when self-hosted. For longer-context needs, Cloudflare points users toward Clef, which keeps its 64k.

Clef gets faster thanks to SGLang

No new weights for Clef: the gains come from the serving layer on Workers AI. For roughly 800 tokens, median latency drops from 262 to 152 ms; for roughly 3,400 tokens, from 616 to 305 ms (2×); for roughly 16,000 tokens, from 2,721 to 1,635 ms.

One of the key optimizations is the switch to SGLang. Cloudflare contributed the Clef integration via PR #42721, expected in SGLang 0.5.22, and publishes launch commands for self-hosting on Hugging Face and in the SGLang cookbooks.

What Cloudflare does with it internally

Cloudflare emphasizes that detection and classification are moving up into the model layer: no more need for a dedicated machine learning team or a custom dataset to cover a specific domain.

Examples cited: the public GitHub repo for documentation automatically closes spam issues, the EmDash CMS moderates plugin libraries against phishing, the data loss prevention team spots personal data like ID documents, and threat intelligence detects malicious domains. Clef is compatible with Jev's API and available via AI Gateway: switching the model identifier is enough to make the switch.

“Instead of setting up cascading pipelines of models that transcribe speech-to-text, or splitting audio and image channels from video, you can just call one model to make decisions across any modality.”
“We decided we wanted to do something in the decision model space on a Friday evening, trained the model over the weekend, and launched it on Thursday.”
“Even a full 21-second video clip with sound is scored in about 1.5 seconds, all in a single API call.”

Why it matters

Cloudflare is carving out a specific and underrated niche: decision models that don't write anything but decide fast, with a calibrated score and a guaranteed format. For agentic workflows, that's often exactly what's needed at every step, far more than a chatty, expensive LLM. Adding audio and video in a single call, under a second and a half, opens up concrete use cases in inspection, moderation, or support. Still, the numbers deserve scrutiny: Clef-omni regresses compared to Clef on most workflow evaluations, and Clef-flash's price cut comes with a context reduction presented as trivial. The weekend-trained-model rhetoric is impressive, but it also owes a lot to the fact that most of the heavy lifting rests on a frozen Qwen backbone and LoRAs. Cloudflare's real strength lies elsewhere: native integration into its edge, its AI Gateway, and its own products, plus compatibility with Jev's API, which makes switching providers trivial. This is as much an infrastructure strategy as a model strategy.

#ai#cloudflare#multimodal#open weights#agents#inference
Original source
Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
Cloudflare
Open the article ↗

For you

Put it to work on your sources.

Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.

For your team

The same machine, on your topics.

A space in your colours, your watch angles, your curators. Pilot open to three companies.

Read next