Utopia Tech
Engineering4 min read

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

Following last week’s release of Clef and Clef-flash , Cloudflare’s open-weight decision models, we decided to bring forth more gifts. Today, we’re releasing Clef-omni, which takes in audio and video input alongside text and image. We also cut the price of Clef-flash so it is now cheaper than Jev, and we made Clef faster. Although the model game is still early for Cloudflare, i

UT

Utopia Tech

October 9, 2026 · 4 min read

Share

Following last week’s release of Clef and Clef-flash , Cloudflare’s open-weight decision models, we decided to bring forth more gifts. Today, we’re releasing Clef-omni, which takes in audio and video input alongside text and image. We also cut the price of Clef-flash so it is now cheaper than Jev, and we made Clef faster.

Although the model game is still early for Cloudflare, innovation and iteration is in our DNA, and we apply these principles to everything we do. In fact, the story of Clef came together over the course of less than a week. We decided we wanted to do something in the decision model space on a Friday evening, trained the model over the weekend, and launched it on Thursday.

Even with such a short timeline, we were able to ship performant, high-quality, open-weight models for the community — imagine what more we can do in the future. For today, we’re excited to keep up the momentum with new additions and improvements to our Clef family of models. This is just the beginning, and we’ll continue to get better, faster, cheaper, and more innovative.

Clef-omni takes audio, video, image, and text input Our new Clef-omni model is able to take audio, video, image, and text input. This changes the paradigm for decision models, which have been largely text-only since the debut of Jev from TypeSafe. With Clef, we supported images and video frame arrays, but Clef-omni is able to take in audio (wav or mp3) and video (mp4 or webm) alongside text and images.

Instead of setting up cascading pipelines of models that transcribe speech-to-text, or splitting audio and image channels from video, you can just call one model to make decisions across any modality. We are now one step closer to a model that is able to interact with the world as we experience it — through audio, visual, and textual communication, all in one.

Check out our developer docs for the new Clef-omni model. We also released the model open-weights on HuggingFace . For a quick start, here’s how you can send new modality inputs to Clef-omni: Clef-omni expands our open-weight decision architecture to natively handle multimodal workflows.

We built this on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts (MoE) foundation, where the model was already capable of processing text, imagery, audio, and video directly within a single pipeline. However, we adopt the primary comprehension backbone while discarding the text-to-speech output components. In production, Clef-omni executes a quick prefill pass across the complete payload, scoring all modalities and valid parameter options simultaneously.

Because we skip output token generation since Clef models are not Large Language Models (LLMs), we remove the overhead of transcribing or captioning incoming files. Media elements map straight into the unified sequence, where video and audio are synced with visual frames for joint processing. Clef-omni then pulls candidate values directly from internal embeddings using our two-stage attention routing: every valid option gathers key evidence from the input (regardless if they are buried in text snippets, visual elements, or audio streams) before field vectors cross-attend across the full context to compute confidence scores.

We add a built-in lexical grammar that preserves option semantics, so that we can deliver fast, schema-constrained scoring across every input type. We train the model with the same techniques as we did with Clef — freezing the Qwen3 backbone, training low-rank adapters (LoRA), and applying our standard post-training approach: combining label-smoothed cross-entropy loss with Brier score calibration.

The result is a model that is resilient and designed to succeed regardless of schema variations, field ordering, and prompt structures. This brings several core upgrades to the Clef model family. Calibrated, schema-bound decisions now seamlessly handle text, images, audio, and synchronized video in a single API call.

It's also fast. Text-only decisions return in about 130 ms at the median, image inputs in about 150 ms, and audio clips in a few hundred milliseconds. Even a full 21-second video clip with sound is scored in about 1.

5 seconds, all in a single API call. Our benchmarking shows strong performance across benchmarks, even with new modalities and MoE architecture. Benchmark Clef-omni Clef Clef-flash Jev BFCL · case exact 98.

2 98. 47 98. 76 95.

75 ToolRet · nDCG@10 66. 6 69. 19 66.

43 65. 28 API-Bank · accuracy 92. 7 91.

93 93. 11 88. 19 Home appliances · case exact 69.

3 82. 95 97. 73 52.

27 When2Call · accuracy 63. 3 72. 37 65.

58 80. 97 BANKING77 · macro-F1 94. 8 94.

20 90. 93 79. 74 CLINC150+OOS · macro-F1 97.

7 97. 43 66. 77 89.

27 BRIGHT · nDCG@10 42. 0 45. 91 39.

26 47. 52 Amazon ESCI · macro-F1 57. 8 57.

48 57. 39 55. 21 PhishNChips · accuracy 73.

2 79. 60 75. 05 62.

55 We also benchmarked against the TypeSafe evals to show how Clef-omni performs . Workflow Metric Clef-Omni Clef Clef-flash Jev Invoice processing Exact actions 60. 2 64.

7 57. 1 61. 8 Invoice processing Primary action 82.

0 86. 2 73. 3 83.

1 Customer service Exact actions 71. 6 76. 3 77.

0 76. 0 Security incidents Exact actions 61. 7 62.

9 61. 7 61. 7 Agent trace observability Primary action 65.

8 68. 5 69. 8 71.

6 Clef-flash model is now cheaper We heard your feedback — you want a decision model affordable enough to incorporate into any workflow. We’ve made a few optimizations and can now offer Clef-flash at cheaper prices. We want Clef to be able to make decisions that scale for you across all your agentic workflows, and now our pricing also incentivizes that.

We were able to optimize the model to the point where Clef-flash is now cheaper than Jev while still being performant and high-quality. Check out the developer docs for the most up-to-date pricing information at all times, including more detail on how image and audio modalities are converted to input tokens for pricing. Clef-flash – originally $0.

09 per M input tokens, now $0. 038 per M input tokens Clef – remains at $0.

Originally published at blog.cloudflare.com

Share
▸ Want a deeper look?

Talk to an architect about applying this to your stack.

60-minute technical evaluation, no obligation. We'll map the ideas in this article to your environment.

Skip to main content