Over the last few weeks, there has been lots of buzz around decision models such as Typesafe AI’s Jev System One model. While classifier models have been around for some time, Jev introduces a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required.
These models are capable enough to work over any set of inputs without constantly retraining the model to incorporate new classification categories. This contrasts with the world of Large Language Models (LLMs), which are largely non-deterministic, but are open-ended enough to reason and generate text and tool calls for agentic workloads. Today, we’re releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI .
Clef is currently the leader when evaluated against the Jev Decision Index , you can view full results on the live benchmark demo site . These models are smarter, faster, and fully Jev-API compatible, so you can experiment with these hosted models easily. We’re fully open-sourcing these models on Hugging Face under an Apache 2.
0 license for you to run locally and experiment with yourselves. Lastly, we’re excited to debut our new reinforcement learning (RL) product, which allows customers to fine-tune Clef to suit their use cases as well. What is a decision model?
A decision model makes classifications to help agents decide how to act, based on certain probabilities. For example, you can pass in a customer support message (inputs) and ask if it is urgent and which team should handle it. A decision model will return typed answers with probabilities (outputs), which your code can use to route the ticket, trigger an escalation, or defer to a human.
This means that a human does not necessarily need to be in the loop for agentic decisions anymore — agents can programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed. Specifically at Cloudflare, we’ve been testing our new Clef model on our Threat Intelligence team to help us classify website domains. By giving a domain to Clef (with Browser Run) it can quickly identify categories that the domain falls under — for example, it might classify a domain with a 95% chance it is a fashion website, 85% ecommerce, <1% phishing, etc.
This classification took our Clef model 2. 2s to fetch, render, and classify the website. In contrast, our fastest general LLM gpt-oss-120b took 4.
7s in the same workflow, and only returned two classifications. As a user, you can imagine how a 2x savings in latency and results can help us improve our threat intelligence workflows and be faster in identifying malicious or legitimate domains. Generalize this to any use case where you need to make quick programmatic decisions, and you unlock powerful agentic workflows that are able to autonomously decide, reason, and execute.
In music theory, a clef is a symbol placed at the beginning of a musical staff that assigns specific pitch names to the lines and spaces. A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it. We chose Clef as the name of our family of decision models, as it serves similar purposes, and the CF hearkens to Cloudflare.
How is Clef different from other decision models? Although the market is getting increasingly saturated with decision models, Clef has some unique properties that make us excited to release it to the public. First, it has a vision encoder so it’s able to take in images and classify visual content.
This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev’s 32k), which allows users to squeeze more input state for the model to classify against. Third, our model is accurate and powerful, scoring competitively against other decision models on the market across various quality benchmarks.
We shortlisted some evaluations below that are important for decision-making as defined by the Jev Decision Index and scored some of the more popular models on the market for it. Check out the table below for benchmarks, or view the scores on our live decision index demo site : Benchmark Clef Clef-flash Jev DiffusionGemma Jev Kev 9B Laya BFCL · case exact 98.
47 98. 76 95. 75 96.
52 94. 51 38. 13 ToolRet · nDCG@10 69.
19 66. 43 65. 28 61.
21 64. 26 12. 69 API-Bank · accuracy 91.
93 93. 11 88. 19 83.
66 56. 30 11. 41 Home appliances · case exact 82.
95 97. 73 52. 27 42.
05 25. 00 0. 00 When2Call · accuracy 72.
37 65. 58 80. 97 75.
44 49. 62 11. 94 BANKING77 · macro-F1 94.
20 90. 93 79. 74 74.
28 84. 83 14. 29 CLINC150+OOS · macro-F1 97.
43 66. 77 89. 27 83.
49 79. 03 3. 19 BRIGHT · nDCG@10 45.
91 39. 26 47. 52 42.
94 38. 53 19. 90 Amazon ESCI · macro-F1 57.
48 57. 39 55. 21 53.
37 49. 22 24. 40 PhishNChips · accuracy 79.
60 75. 05 62. 55 85.
35 50. 75 50. 15 We also ran benchmarks across Typesafe’s own eval suite and our Clef models fared well, beating Jev in 3 out of 4 areas.
Notably, our Clef-flash performs exceptionally well, given how much faster it is. Workflow Clef Clef-flash Jev Invoice processing 64. 7 57.
1 61. 8 Customer service 76. 3 77 76.
0 Security incidents 62. 9 61. 7 61.
7 Agent trace observability 68. 5 69. 8 71.
6 Across the 43 eval benchmarks that we ran, our Clef models beat the decision models on latency (except for Laya which is very fast but trades off quality in the benchmarks above): Benchmark Clef Clef-flash Jev DiffusionGemma Jev Kev-9B Laya Median latency · ms 209. 3 38. 8 524.
1 84. 4 51. 4 5.
8 p95 latency · ms 238. 6 122. 4 536.
0 211. 2 187. 9 222.
5 On top of the latency benefits from the model itself, our Clef models are hosted on Workers AI. Because they are hosted on Cloudflare’s infrastructure, we’re able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions.
Originally published at blog.cloudflare.com

