Opens in a new tab

Cloudflare releases Clef and Clef-flash

02-10-2026

Cloudflare released Clef and Clef-flash on 1 October 2026, two open-weight decision models that return probabilities instead of text. Clef-flash decides in 38.8 milliseconds at the median, both carry a 64,000 token context and both read images and video.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

cloudflare releases clef and clef flash, open weight decision models at 38.8ms
Sign up for our Newsletter

Cloudflare released Clef and Clef-flash on 1 October 2026, two open-weight models that answer a fixed set of questions with probabilities instead of text. Clef-flash returns a decision in 38.8 milliseconds at the median. Both ship under Apache 2.0 on Hugging Face and run hosted on Cloudflare’s Workers AI.

What are Cloudflare Clef and Clef-flash?

A decision model does one job. You hand it a document and a list of allowed answers, and it returns a probability for each one. It never writes a sentence. That fits the part of an AI agent that is not writing but choosing: which tool to call, whether an input is safe, which of twelve support categories a ticket belongs in.

Clef has 27 billion parameters and is built on Qwen3.8-27B. Clef-flash has 9 billion and is built on Qwen3.5-9B. Both handle three question shapes: yes or no, pick one named option, and rate on a scale. Both take up to 64 questions in a single request. Both read images and video frames alongside text, which the incumbent decision model, TypeSafe’s Jev, does not.

Cloudflare also kept the Jev API. Clef answers the same calls, so swapping is a change of endpoint rather than a rewrite.

Clef benchmarks, latency and context window

The figures worth comparing are speed, context length and the one benchmark where Clef’s lead is wide.

AspectClefCompetitor
Median decision latency209.3 msClef-flash: 38.8 ms
Context window64,000 tokensJev: 32,000 tokens
BANKING77 (sorting bank support messages)94.20 macro-F1Jev: 79.74
BFCL (picking the right tool to call)98.5%Clef-flash: 98.76% exact match
Image and video inputYesJev: text only
Licence and weightsApache 2.0 on Hugging FaceJev: hosted API only

These are Cloudflare’s own numbers, measured across 43 evaluation sets and published with the model card. Cloudflare does not say where the Jev figures came from, so read the head to head rows as the vendor’s framing rather than an independent rerun.

What Cloudflare is not saying about Clef

Two things stand out. The announcement carries no price. The changelog points at a pricing page instead, and the figures now in circulation, $0.24 per million input tokens for Clef and $0.09 for Clef-flash, come from secondary coverage rather than from Cloudflare. For a model sold on cost per decision, that is the number you would expect first.

The second is the benchmark Cloudflare does not lead with. Jev still beats Clef on GPQA Diamond and MMLU-Pro, the two sets that test knowledge rather than sorting. If your decisions need the model to know things, rather than to file things it has been handed, the gap runs the other way.

How Clef fits the decision model pile-up

Four decision models landed inside eight days. Liquid AI shipped d1 on 30 September. Cloudflare and Amazon both shipped on 1 October, Amazon with Strands Decider 2B. All of them point at Jev, and all of them make the same argument: an agent that asks a 27B model to write JSON is paying for a capability it does not use.

Cloudflare’s angle is distribution rather than architecture. Clef runs on the same edge network your Workers already sit on, which is why 38.8 milliseconds is plausible in production and not only on a benchmark rig. The company also announced a reinforcement learning service for fine-tuning Clef on your own traffic, built from AI Gateway, Workers AI and Containers. That part is not self-serve. It starts as a hands-on engagement with Cloudflare engineers, with a public platform promised later.

What this means

Worth testing now if you already run Workers and your agent makes more routing calls than generation calls. Clef-flash is the model to try first, not Clef. At 38.8 milliseconds a decision stops being a step the user waits through, and a 9B model at that speed is cheap enough to put in front of every request rather than the uncertain ones. The vision input is the genuine differentiator here: screenshot triage, document type detection and moderation of uploaded images all become a single call instead of a pipeline.

Safe to ignore if you are not on Cloudflare and you do not want to host 27 billion parameters yourself. The Apache 2.0 licence is real and the weights are there, but self-hosting Clef buys you a classifier, not a general model, and a fine-tuned small encoder will often match it on one narrow task for a fraction of the hardware. The thing to watch is not the benchmark lead over Jev, which is one vendor’s table. It is whether four labs shipping the same product in eight days means decision models become a commodity layer by Christmas, in which case the price Cloudflare has not published is the only number that will matter.

For more information, visit the official announcement of Clef on the Cloudflare blog.

Add DataNorth AI to your Google favorites