Clef: Cloudflare Launches Its Own Open Source Decision Models
With Clef and Clef-flash, Cloudflare enters the nascent market of "decision models" and adds a reinforcement learning fine-tuning platform to the mix.

Who covers it
Picked up by 3 outlets · 1 tech · 2 forums · 4 social posts
In brief
Cloudflare has released Clef and Clef-flash, two open source decision models under the Apache 2.0 license, hosted on Workers AI and compatible with the API of Jev, the model from Typesafe AI that launched the category. Faster than Jev according to Cloudflare's own benchmarks, they accept images and a 64k token context. Cloudflare is also launching an RL fine-tuning service, first white-glove, then self-serve.
🍺 Bar-stool version
An LLM is the brilliant coworker you ask "urgent or not?" who hands you back three nuanced paragraphs. A decision model is the one who answers "yes, 92% confidence," period, in a fifth of a second. Cloudflare just released its own, free, API-compatible with the competitor's, and then offers to train it on your own data, on its servers, naturally. The day agents start making decisions on their own, the real power goes to whoever supplies the model that does the deciding.
Key takeaways
- 1
Clef (based on Qwen3.8-27B) and Clef-flash (Qwen3.5-9B) are published on Hugging Face under the Apache 2.0 license and hosted on Workers AI.
- 2
Both models are fully compatible with the API of Jev System One, Typesafe AI's decision model, allowing you to switch between them without rewriting code.
- 3
Clef adds a vision encoder to classify images, where Jev only handles text, and doubles the context window (64k vs 32k).
- 4
In median latency, Clef responds in 209.3 ms and Clef-flash in 38.8 ms, versus 524.1 ms for Jev, across 43 benchmarks run by Cloudflare.
- 5
On Typesafe's eval suite, Clef beats Jev in 3 out of 4 domains (invoices, customer service, security incidents) but loses on agent trace observability.
- 6
The decision step is non-autoregressive: a single prefill pass on Qwen, then parallel scoring of the schema's valid choices, with no intermediate text generation.
- 7
Cloudflare is launching a reinforcement learning fine-tuning service, first with its forward-deployed engineers team, then as a self-serve platform.
A new category: the decision model
Cloudflare starts from an observation: over the past few weeks, Jev System One, Typesafe AI's model, has popularized the idea of a "decision model." It's not an LLM that reasons and generates text, but a model that produces structured, bounded outputs, quickly, cheaply, and consistently.
The principle: you provide a state (a support ticket, for example) and typed questions. The model returns answers with probabilities: is it urgent, which team should handle it, what's the severity. The code can then route, escalate, or hand off to a human.
Unlike classic classifiers, these models don't need to be retrained every time a new category appears: the options are defined in the request itself. Cloudflare sees this as a building block for agents that decide and act without a human in the loop, except when they judge it necessary.
The internal use case: classifying domains
Cloudflare's Threat Intelligence team tested Clef, paired with Browser Run, to categorize domain names. Example given: 95% fashion, 85% e-commerce, under 1% phishing.
The whole pipeline (fetching, rendering, and classifying the page) takes 2.2 seconds with Clef. With gpt-oss-120b, Cloudflare's fastest general-purpose LLM, the same workflow takes 4.7 seconds and returns only two categories.
The name itself comes from music: the clef placed at the start of a staff fixes the names of the notes that follow, just as the model fixes the framework for upcoming actions. And "CF" nods to Cloudflare.
What Clef brings compared to Jev
Three arguments are put forward. First, vision: Clef can classify visual content. Second, context: 64k tokens versus 32k for Jev. Finally, quality, measured on benchmarks drawn from the Jev Decision Index.
The gaps are sometimes striking. On BANKING77, Clef scores 94.20 macro-F1 versus 79.74 for Jev; on CLINC150+OOS, 97.43 versus 89.27; on the "home appliances" task, Clef-flash reaches 97.73 versus 52.27 for Jev.
But Jev stays ahead on When2Call (80.97 vs 72.37) and BRIGHT (47.52 vs 45.91). Clef-flash also collapses on CLINC150+OOS (66.77). On PhishNChips, another model in the table beats Clef (85.35 vs 79.60).
On latency, Clef-flash posts 38.8 ms median and 122.4 ms p95. Only Laya is faster (5.8 ms median), but its quality scores are very low. Edge GPU hosting also cuts network latency, which lets you place Clef in an agent's hot path.
Under the hood: no generation, just scoring
Clef builds on an experiment Cloudflare published the same week Jev launched: adapting DiffusionGemma to extract deterministic probabilities via logprobs, drawing on work by Matt Mastracci, a vLLM contributor.
The backbone is now Qwen, frozen: Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash. Cloudflare trains a routing head and rank-256 LoRA adapters. At inference time, a single prefill pass is enough, then the schema's valid choices are scored in parallel, with no token-by-token text generation.
The architecture relies on a two-stage attention routing: each option retrieves the relevant parts of the prompt, then the fields attend to each other and re-read the payload before scoring. Training combines cross-entropy with label smoothing and a Brier loss to calibrate probabilities, on synthetic data that permutes field order, prompts, and schemas.
Added to this is RLCD (Reinforcement Learning for Calibrated Decisions): partial credit for nearby ordinal choices, reward for fully correct records, and a reference penalty to prevent distribution drift.
The real bet: RL fine-tuning
Cloudflare's internal teams want specialized versions of Clef: triaging Trust & Safety reports, sorting support requests, distinguishing good from bad crawlers in Bot products. Cloudflare says it has over 15 years of network data and labeled decisions to draw on.
The service starts with the FDE (forward-deployed engineers) team, in a white-glove mode. The goal is then a self-serve platform to capture data, fine-tune, and redeploy, entirely within Cloudflare.
The pipeline relies on existing building blocks: AI Gateway to automatically build datasets from traffic, Workers AI to generate rollouts, Containers as an RL sandbox, a new Trainer component to update weights, and Bring Your Own Model (Cog, inherited from the Replicate acquisition) to redeploy.
“A decision model makes classifications to help agents decide how to act, based on certain probabilities.”
“The decision step is non-autoregressive, so there's no intermediate text to generate token by token.”
“We believe that Clef has the ability to disrupt the way we use agents, which fits naturally into Cloudflare's mission of being the agent cloud.”
Why it matters
This announcement shows how quickly a category can become commoditized: just weeks after Jev appeared, Cloudflare released an open source equivalent, API-compatible, and faster. For Typesafe AI, this is a direct threat: when a competitor copies your interface and gives away the weights, all you have left is quality and trust. The deeper stakes lie elsewhere: decision models are the missing link that makes autonomous agents possible, and Cloudflare wants to be "the agent cloud" that hosts, trains, and redeploys them. Open source here serves as a loss leader; the captured value will come from fine-tuning and edge hosting, which create strong platform lock-in. The numbers should also be read with caution: the benchmarks are selected and run by Cloudflare, Jev still has the edge on several tasks, and Clef-flash shows sharp drops. Finally, the idea that a human is no longer strictly necessary in the loop assumes well-calibrated probabilities; that's precisely what Cloudflare claims to be optimizing for, but it's in production that this will truly be tested.
Free account
You just read an AI Sources article
Create a free account: a month of archives in full, your own sources summed up like this one, your notes and highlights.
For you
Put it to work on your sources.
Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.
For your team
The same machine, on your topics.
A space in your colours, your watch angles, your curators. Pilot open to three companies.
Read next
#geminiYesterdayGemini 4 Argon: Google Releases Its Most Powerful Model, But Defenders Get It First
Google announces a frontier model that rewrites kernels in Rust and hunts down vulnerabilities, then locks it away while it tightens the guardrails.
Source · Le blog de Google (The Keyword) · Gemini 4 Argon: our next era of frontier intelligence
#evaluationYesterdaySteering a LLM has a price: Apple measures what control costs in fluency
A study from Apple and Pompeu Fabra University shows that the most effective control methods are often the ones that damage generated text the most.
Source · Apple Machine Learning Research · On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study
#californiaYesterdayCalifornia: Newsom Signs 13 AI Laws, From HR to Gene-Synthesis Labs
Algorithmic layoffs, workplace surveillance, deepfakes, synthetic DNA: Sacramento keeps stacking up safeguards while Washington looks elsewhere.
Source · Governor of California · California's nation-leading AI framework just got stronger, Governor Newsom signs more first-in-the-nation worker protections and more