Cloudflare Releases Clef and Clef-flash, Open-Weight Decision Models on Workers AI
Cloudflare launched Clef and Clef-flash on 1 Oct 2026: Workers AI decision models with Apache 2.0 weights, USD list pricing, and vendor-reported latency wins over Jev.
Cloudflare released Clef and Clef-flash on 1 October 2026, the first models trained by its Workers AI team. Both are decision models: they take a state plus typed questions and return probabilities, not free-form text.
The company hosts them on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, open-sourced the weights under Apache 2.0 on Hugging Face, and says the API is drop-in compatible with TypeSafe AI’s Jev System One format. Official list prices are USD $0.24 per million input tokens for Clef and $0.09 for Clef-flash.
Benchmark wins and latency figures in this piece are Cloudflare-reported unless noted. Independent Decision Index ranking of the self-reported scores had not been confirmed in secondary coverage at publish time.
What a decision model actually does
Unlike a chat LLM, Clef does not write an answer. You send state (text, JSON, and optionally images or video) and up to 64 typed questions. Each question is one of three shapes:
- noul: yes/no, returns the probability of yes
- choice: pick one option from a set you define, with per-option probabilities
- score: rate against an ordered rubric, with a probability-weighted score
Your code then routes a ticket, blocks a request, or escalates. Cloudflare’s launch post frames this as the hot path for agents that need a cheap, bounded decision before a larger model acts.
The company says its Threat Intelligence team used Clef with Browser Run to fetch, render, and classify a domain in 2.2 seconds, versus 4.7 seconds for gpt-oss-120b in the same workflow. That comparison is vendor-run, on Cloudflare’s own stack.
Clef vs Clef-flash: size, price, context
| Model | Size (vendor) | Workers AI ID | List price (USD) | Context |
|---|---|---|---|---|
| Clef | 27B (Qwen3.8-27B backbone) | @cf/cloudflare/clef | $0.24 / 1M input tokens | 65,536 tokens |
| Clef-flash | 9B (Qwen3.5-9B backbone) | @cf/cloudflare/clef-flash | $0.09 / 1M input tokens | 65,536 tokens |
Pricing and context come from Cloudflare’s Clef and Clef-flash model pages. Both support vision. Docs allow up to four embedded PNG, JPEG, or WebP images per request (size caps apply). Remote image URLs are not accepted.
Michelle Chen, Cloudflare AI Platform group product manager, told The Register that local runs need about 85 GB VRAM for Clef and 41 GB for Clef-flash at single concurrency with a 64k context. Training datasets are not public, even though weights are Apache 2.0.
What Cloudflare claims against Jev
Cloudflare positions Clef against TypeSafe’s Jev: same System One question types, plus vision and a 64k context window. The Register notes Jev can also handle up to 64k tokens across a request, with a 32k limit on state plus the longest individual question.
Vendor-reported median latency across 43 runs: Clef 209.3 ms, Clef-flash 38.8 ms, Jev 524.1 ms. Cloudflare says that puts Clef about 2.5x faster than Jev at the median, and Clef-flash about 13x faster.
On Cloudflare’s published Decision Index-style table, Clef or Clef-flash beats Jev on 8 of 10 benchmarks. Examples Cloudflare highlights:
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL (case exact) | 98.47 | 98.76 | 95.75 |
| BANKING77 (macro-F1) | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS (macro-F1) | 97.43 | 66.77 | 89.27 |
| Home appliances (case exact) | 82.95 | 97.73 | 52.27 |
On TypeSafe’s own workflow evals, Cloudflare says Clef beats Jev on invoice processing, customer service, and security incidents, and trails on agent trace observability. Treat all of these as vendor-reported until third parties reproduce them on the live Decision Index.
The Register also notes Cloudflare’s hosted Clef list price ($0.24/M) is several times TypeSafe’s Jev rate it cited ($0.042/M). Open weights are the cost escape hatch if you have the GPUs.
How Clef was trained (vendor description)
Cloudflare says both models freeze a Qwen backbone, run a prefill-only pass, then score valid schema choices in parallel without autoregressive text. Post-training used label-smoothed cross-entropy, Brier loss for calibration, and something it calls Reinforcement Learning for Calibrated Decisions (RLCD).
That architecture story is Cloudflare’s. What developers can verify today is the hosted API, the Hugging Face weights, and whether the typed outputs match the System One schema in their own harness.
Fine-tuning and the RL product
Alongside the models, Cloudflare announced a reinforcement learning fine-tuning service. Near term it is hands-on with a forward-deployed engineer team. Longer term it wants a self-serve loop that uses AI Gateway traffic as training data, Containers as RL sandboxes, a new Trainer to update weights, and Bring Your Own Model redeploy on Workers AI.
Internal use cases Cloudflare named include Trust and Safety scoring, support triage, and bot good/bad decisions. Those are product pitches, not customer case studies.
Confirmed vs still vendor-only
| Claim | Status |
|---|---|
| Clef and Clef-flash live on Workers AI (1 Oct 2026) | Confirmed (Cloudflare blog and changelog) |
| Apache 2.0 weights on Hugging Face | Confirmed (Cloudflare) |
| USD $0.24 / $0.09 per 1M input tokens | Confirmed (Workers AI model pages) |
| Jev-compatible System One API | Confirmed as Cloudflare design claim; verify in your stack |
| Vision (up to 4 images) and video state | Docs-confirmed for images; video listed on model pages |
| Beats Jev on latency and most listed benchmarks | Vendor-reported |
| 2.2s vs 4.7s domain classification vs gpt-oss-120b | Vendor-reported internal Threat Intel test |
| 85 GB / 41 GB VRAM for local Clef / Clef-flash | Reported by Michelle Chen to The Register |
| Training data released | No (Chen to The Register) |
What it means for developers in the US, Canada, Australia, and India
Workers AI is a global Cloudflare product. If you already have an account, Clef is available through the same env.AI.run() binding and REST path as other Workers AI models, and through AI Gateway. There is no separate US-only or India-only gate described in the launch materials.
Practical checklist:
- If you already speak Jev’s System One schema, point the endpoint at Clef and A/B the probabilities on your own labels before you trust routing.
- Price the hot path: $0.09/M for Clef-flash may still beat a large LLM call for triage, but it is not free, and The Register’s Jev price comparison matters if cost is the only reason to switch.
- Local open weights help teams that cannot send state off-network, provided they can spare roughly 41 GB or 85 GB VRAM.
- Pair decisions with the rest of Cloudflare’s agent stack you already know, including agent wallets for spend caps, rather than treating Clef as a full agent by itself.
Agent infrastructure keeps widening. Related context on this site includes Supabase betting on per-agent SQLite and NVIDIA OpenShell for agent sandboxes. Clef is the classification slice of that stack, not the whole stack.
Frequently Asked Questions
What did Cloudflare release on 1 October 2026?
Two decision models, Clef (27B) and Clef-flash (9B), hosted on Workers AI, with Apache 2.0 weights on Hugging Face, plus an early RL fine-tuning service for design partners.
How much do Clef and Clef-flash cost?
Official Workers AI list prices are $0.24 per million input tokens for Clef and $0.09 per million input tokens for Clef-flash. Local Hugging Face weights avoid that bill if you run your own GPUs.
Is Clef better than Jev?
Cloudflare reports lower median latency and higher scores on most of the Decision Index-style benchmarks it published, and wins on 3 of 4 TypeSafe workflow evals. Those results are vendor-reported. Run your own labels before you cut over.
Can Clef see images?
Yes. Cloudflare’s docs say both models are multimodal and accept up to four embedded images per request. Model pages also list video as accepted state. Jev, as described in Cloudflare’s post, is text-only today.
Are the training datasets open?
No. Weights are Apache 2.0. Michelle Chen told The Register the training datasets are not public.