Cloudflare Clef: open-weight models that return decisions
Cloudflare's Clef (27B) and Clef-flash (9B) answer typed questions instead of writing text, at 209 ms and 39 ms median latency, under Apache 2.0.
3 min read

By the numbers
- parameters in Clef, built on Qwen 3.8-27B
- 27B
- parameters in Clef-flash, built on Qwen 3.5-9B
- 9B
- Clef-flash median latency
- 38.8 ms
- token context window
- 64k
- Clef-flash
- 38.8 ms
- Clef
- 209.3 ms
- Jev
- 524.1 ms
Cloudflare released two open-weight decision models, Clef and Clef-flash, on October 1, 2026. They do not write text. They read an input and answer a fixed set of typed questions with probabilities, which lets an AI agent pick its next step quickly. Both are free to download under the Apache 2.0 license and run on Cloudflare's hosted Workers AI service, according to Cloudflare's announcement.
What a decision model is
"A decision model makes classifications to help agents decide how to act, based on certain probabilities," Cloudflare writes. A chat model produces an answer one token at a time. Clef skips that step. Cloudflare calls the design non-autoregressive, meaning it does not generate output word by word.
Traictory explains the mechanism. Each model reads the whole input in one pass, then scores every possible answer in parallel. Nothing is generated, so there are no reasoning tokens to wait for.
This is the same idea behind TypeSafe's Jev, a closed model launched on September 15, and Convai's Laya, a 421-million-parameter open model released four days later. Clef brings the idea to a hosted platform that many developers already use.
The two models
| Clef | Clef-flash | |
|---|---|---|
| Parameters | 27 billion | 9 billion |
| Base model | Qwen 3.8-27B | Qwen 3.5-9B |
| Median latency | 209.3 ms | 38.8 ms |
| Context window | 64k tokens | 64k tokens |
| License | Apache 2.0 | Apache 2.0 |
Cloudflare reports a 95th-percentile latency of 122.4 ms for Clef-flash. That means 95% of its requests finished within that time.
Both models are built on Qwen, the open model family from Alibaba. Traictory reports that the Qwen base stays frozen. Cloudflare trained small add-on layers, called rank-256 low-rank adapters, on top of it.
Clef also has a vision encoder, so it can take images as input. Cloudflare presents that as a difference from Jev, which handles text only.
Speed and price claims
Cloudflare says both models cut latency in half compared with gpt-oss-120b, an open OpenAI model, in its own threat intelligence tests. It says it evaluated the models on 43 benchmarks.
The price drew attention. Traictory puts Clef at $0.24 per million tokens on Workers AI, against $0.042 per million for Jev. That makes Clef nearly six times as expensive per token. Traictory lists a median latency of 524.1 ms for Jev, against 209.3 ms for Clef.
For hosted use, Cloudflare says it does "not read, store, or train on your requests." The weights are also on Hugging Face, so a team can run them on its own hardware and avoid the per-token price entirely. The same announcement introduces a new reinforcement learning fine-tuning platform for training such models.
The open questions
Traictory reports that Hacker News commenters questioned both the price and the benchmark method. "Public benchmarks are easy to cheat," one commenter wrote. Traictory also notes that "no third-party run is cited" for the quality claims.
It lists three gaps: independent validation, tests against the current version of Jev, and data from real production use. None of those exist yet.
What this means for developers
Look at the decisions your agent makes before every step. Routing a request, flagging a message or choosing a tool are all classification jobs. If a large chat model handles them today, a decision model could cut that step from seconds to milliseconds.
Start with Clef-flash. At a 38.8 ms median, it is fast enough to sit in the request path of a live app. Move to the 27B model only if your own tests show the smaller one gets too many answers wrong.
Run your own evaluation before you switch. The published numbers come from Cloudflare, and no outside group has reproduced them. Take 200 to 500 real inputs from your logs, label the correct answers, and compare Clef with what you use now.
Compare costs both ways. On Workers AI you pay per token. With the Apache 2.0 weights, you pay for your own GPUs instead. A busy classification step can make self-hosting the cheaper choice, and it keeps the data in-house.
Sources
Related articles

Strands Decider 2B makes agent choices in 115 ms
AWS's Strands Decider 2B is an open-source model that picks from a list instead of writing text. It answers in 115 ms on an RTX 3090 and scores 72.3% on JevBench.

AstaBrief 8B: Ai2's open model writes cited reports
Ai2 released AstaBrief 8B, an Apache 2.0 model on Qwen3-8B that writes cited research reports in 51.1 seconds, 3.5x faster than Asta's Thinking mode.

Qwen's research license, explained
The Qwen Research License lets anyone download Qwen-Image-2.1 but limits it to research or evaluation. Commercial use needs a separate license from Qwen.
The weekly digest
One email every Friday with the week's top stories from all six desks: AI, dev tools, coding, the tech industry, startups and what's next. Free, no spam.