Blog Article
Cloudflare Clef vs Jev: What Decision Models Mean for Your AI Digital Transformation
Cloudflare launched Clef, an open-weight decision model that takes on Jev. Benchmarks, price, and where decision models fit in an AI digital transformation.
Two weeks after Jev put decision models on the map, Cloudflare has shipped its own, and it changes how you can plan an AI digital transformation.
On October 1, Cloudflare released Clef and Clef-flash, two decision models hosted on Workers AI, with the weights open-sourced on Hugging Face under Apache 2.0. Cloudflare says Clef leads on most public decision benchmarks and is faster than Jev. This post covers what the numbers say, where Jev still wins, what each model costs, and where a decision model fits when you modernise a business with AI.
If decision models are new to you, start with our comparison of Jev, Laya and Kev. This post picks up from there.
What Is Cloudflare Clef?
Clef is a decision model. It does not write text. You give it a state (a support message, a domain, an invoice) and a set of typed questions. It returns a probability for every allowed answer, so your code can route a ticket, block a request or hand off to a human.
Cloudflare’s changelog describes it as being in the same family as Typesafe’s Jev. There are two sizes:
- Clef is 27 billion parameters, built on Qwen3.8-27B, with a 64k-token context window.
- Clef-flash is 9 billion parameters, built on Qwen3.5-9B, aimed at latency-critical paths.
Both models have a vision encoder and can take up to four images alongside the state. Cloudflare notes that Jev only classifies text today. The API is Jev-compatible, so Cloudflare says teams can swap models without rewriting their integration.
Clef vs Jev: What Do the Benchmarks Say?
These are Cloudflare’s own numbers. The Register notes that Cloudflare self-reported its scores and that they have not yet been reproduced for ranking on the official Jev Decision Index. Treat them as a strong claim, not a verdict.
Across 10 decision benchmarks, Cloudflare says a Clef model scores highest on 7. A few results from its published table:
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL (case exact) | 98.47 | 98.76 | 95.75 |
| BANKING77 (macro-F1) | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS (macro-F1) | 97.43 | 66.77 | 89.27 |
| When2Call (accuracy) | 72.37 | 65.58 | 80.97 |
| BRIGHT (nDCG@10) | 45.91 | 39.26 | 47.52 |
Jev still comes out ahead on When2Call and BRIGHT. On Typesafe’s own workflow evals, Clef beat Jev in three of four areas: invoice processing, customer service and security incidents. It lost on agent trace observability, 68.5 against 71.6.
Clef-flash also has a weak spot. On CLINC150+OOS it scores 66.77, far below Clef and Jev. The small model is not a free upgrade on every task.
How Fast Is Clef?
Cloudflare’s published latency, across its 43 benchmark runs:
| Latency | Clef | Clef-flash | Jev |
|---|---|---|---|
| Median | 209.3 ms | 38.8 ms | 524.1 ms |
| p95 | 238.6 ms | 122.4 ms | 536.0 ms |
Cloudflare also gave a real workload. Its threat intelligence team uses Clef with a browser tool to classify website domains. Clef took 2.2 seconds to fetch, render and classify a page and returned several categories with probabilities. Its fastest general LLM, gpt-oss-120b, took 4.7 seconds and returned two.
The speed comes from how Clef works. It runs one pass over the input and scores all the allowed answers in parallel, so it never generates text token by token.
Clef Price and Hardware
Per The Register, Clef costs $0.24 per million tokens against $0.042 for Jev, nearly six times as much. Crypto Briefing reports $0.09 per million input tokens for Clef-flash. We could not confirm these rates on Cloudflare’s own pricing page, so check it before you budget.
You can also skip the hosted price and run the weights yourself. Cloudflare’s Michelle Chen told The Register that Clef-flash needs a GPU with at least 41 GB of VRAM and Clef needs 85 GB, assuming one request at a time and a 64k context window.
One caveat on “open source”. The weights are Apache 2.0, but Chen confirmed to The Register that the training datasets are not public.
Cloudflare is also launching a reinforcement learning fine-tuning service for Clef, starting with hands-on help from its engineers for design partners and a self-serve platform later.
Where Decision Models Fit in an AI Digital Transformation
Most AI digital transformation work is not about one clever model. It is about taking a process that people run by hand, such as triaging requests, checking invoices or routing leads, and putting software in charge of the routine decisions while people keep the hard ones.
That is the job decision models are built for. A large language model reads and writes. A decision model makes the bounded call in the middle: urgent or not, which team, which category, escalate or not. Cloudflare’s own example is the same shape: a support message goes in, and typed answers with probabilities come out that your code can use to route the ticket, escalate or defer to a human.
This is our view, not Cloudflare’s: the step that makes a transformation hold up is not the model choice. It is the work around it, which is picking the right process, connecting it to your existing systems, setting the confidence thresholds, and logging every decision. Models will keep changing. The workflow around them is what you keep.
What Enterprises Should Test Before Switching
Everything below is our advice, not Cloudflare’s.
- Use your own labelled data. Take 500 to 1,000 real tickets, invoices or requests with the right answer already known. Public benchmarks do not tell you how a model handles your categories.
- Test the mix, not the average. Clef-flash does very well on some benchmarks and poorly on others. Check the cases that cost you money when they go wrong.
- Set a confidence threshold. The value of a decision model is the probability it returns. Route low-confidence cases to a person, and measure how often that happens.
- Price per decision, not per token. Compare hosted cost against the GPU cost of running the weights yourself at your real volume.
- Keep the swap cheap. Both Clef and Jev use the same API shape, so keep your calls behind one thin wrapper and you can change models when the next release lands.
- Log every decision. A model that decides without a record is hard to defend in an audit. This matters more if you work under the EU AI Act, which we cover in why EU AI Act compliance is an engineering problem.
For where a decision model sits next to the agents it controls, see OpenAI Dots vs Meta Muse. Meta built a second agent just to monitor Muse’s planned actions, which is the same idea: keep the decision step separate from the agent that acts.
Frequently Asked Questions
What is Cloudflare Clef?
Clef is an open-weight decision model that Cloudflare released on October 1, 2026. It answers typed questions about an input and returns a probability for each allowed answer, instead of generating text. It comes in two sizes, Clef and Clef-flash, and runs on Cloudflare Workers AI or on your own GPU.
Is Clef better than Jev?
On Cloudflare’s own benchmarks, Clef scores highest on 7 of 10 tests and is faster. Jev still leads on When2Call and BRIGHT, and on Typesafe’s agent trace observability eval. Cloudflare’s scores have not been independently reproduced yet, so test both on your own data.
How much does Clef cost?
The Register reports $0.24 per million tokens for Clef, against $0.042 for Jev. Crypto Briefing reports $0.09 per million input tokens for Clef-flash. Check Cloudflare’s pricing page for current rates.
Can I run Clef locally?
Yes. The weights are on Hugging Face under Apache 2.0. Per Cloudflare, Clef-flash needs a GPU with at least 41 GB of VRAM and Clef needs 85 GB, for a single request at a 64k context window.
Is Clef a drop-in replacement for Jev?
Cloudflare says the API is fully Jev-compatible, so you can swap models easily. Still run your own tests first, because accuracy differs by task.
The Bottom Line
Clef shows that decision models are not one company’s product. Anyone can now run a fast, bounded decision step inside an agent, and the choice comes down to accuracy on your data, latency and cost per decision. Read the vendor tables as a starting point and run your own test. The lasting value is in the workflow you build around the model.
How Incresco Helps
Incresco does AI digital transformation. Our AI transformation work covers the pieces described above:
- AI strategy and consulting: roadmaps, readiness assessments and the business case, so you pick the right process first.
- Custom AI development: AI agents, copilots and intelligent automation built around your workflows.
- AI integration: connecting AI to your existing systems through workflow automation, API integration and legacy system modernisation.
If you want to test a decision model on one of your own workflows, or plan where agents fit in your operations, talk to our team.