Cloudflare Clef: Open-Source AI Decision Models Explained | Panda Prompt
CREATIVE WORKFLOWS
Cloudflare Open-Sources Clef: Decision Models That Answer in Milliseconds
On October 1, 2026, Cloudflare released Clef and Clef-flash — open-weight decision models that answer typed questions with probabilities instead of generating text, built for AI agents that must classify, route, and act in milliseconds. Here is what decision models are, what Clef changes, and how to think about using them.
EDITORIAL GUIDE2 min read
What Cloudflare announced
On October 1, 2026, during its Birthday Week, Cloudflare released Clef and Clef-flash — the first models trained by its own Workers AI team — alongside a reinforcement-learning fine-tuning platform. The models are decision models: instead of generating text, they answer a fixed set of typed questions (yes/no, pick-one, score) with a calibrated probability for every allowed answer, in a single forward pass.
The weights are published on Hugging Face as Cloudflare/clef and Cloudflare/clef-flash under the Apache 2.0 license, and both models are hosted on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. The announcement blog post is titled "Introducing Clef: our open-source decision models, and new RL fine-tuning platform."
The spec sheet
Clef: 27 billion parameters, built on Qwen 3.8, tuned for precision
Clef-flash: 9 billion parameters, built on Qwen 3.5, tuned for latency
64K context window
Vision encoder — reads text, JSON, images, and video (rival Jev is text-only)
Jev-API compatible — code written for TypeSafe's Jev largely works unchanged
Apache 2.0 license — download, modify, and ship commercially
What "decision model" actually means
Advertisement
Ad Placement (guide-detail-inline)
A normal program decides with an if-statement — but only when the condition is computable. "Is this support ticket urgent?" or "which team owns this incident?" needs understanding, not arithmetic. The usual answer is to ask a chat model and parse its prose. A decision model skips the prose: you define the possible answers up front, the model returns a probability for each, and your code branches on the numbers. One forward pass, no token-by-token generation.
The benchmarks Cloudflare claims
Cloudflare reports a median response of 209.3 ms for Clef versus 524.1 ms for TypeSafe's Jev, and 38.8 ms for Clef-flash. On quality, Cloudflare lists 94.20 macro-F1 on BANKING77 (Jev: 79.74) and 98.47 on BFCL case-exact (Jev: 95.75). One important caveat: these numbers are Cloudflare's own, and Jev still wins several harder reasoning tests in Cloudflare's own tables. Treat them as directionally interesting, not settled.
The RL fine-tuning platform
Alongside the weights, Cloudflare launched a reinforcement-learning fine-tuning service for customizing Clef on customer data. It starts as a hands-on engagement with Cloudflare's forward-deployed engineer team; the planned self-serve path stitches together AI Gateway, Workers AI, and Containers so customers can capture data, fine-tune, and redeploy on Cloudflare's network.
Why this matters for AI agents
Decision models are becoming the hot path inside agents: routing, moderation, validation, next-action selection. Cloudflare is one of two challengers to the closed decision-model category to arrive in a single day — Amazon shipped Strands Decider 2B the same day, an open-source Jev rival for local CPU or GPU. With weights under Apache 2.0, a Jev-compatible API, and llama.cpp already adding a /v1/systemone endpoint for decision models, the ecosystem for fast, local, inspectable decisions is forming quickly.
“A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it.”
Frequently Asked Questions
What is a decision model?
A model that answers a fixed set of typed questions — yes/no, pick one, rate on a scale — with a probability for each allowed answer, instead of generating text. It returns the result in a single forward pass, which makes it fast and easy to wire into code that branches on the numbers.
Are Clef's weights really open source?
Yes — Cloudflare published both models on Hugging Face under the Apache 2.0 license, which permits using, modifying, and shipping the weights commercially under the license terms. Note that the training data is not public, and the hosted Workers AI path has its own pricing.
How is Clef different from TypeSafe's Jev?
Both are decision models with compatible APIs, so most Jev code works with Clef unchanged. The differences: Clef reads images and video (Jev is text-only), Cloudflare claims lower latency and better scores on several benchmarks, and Clef is open-weight while Jev is closed. Jev still wins some harder reasoning tests in Cloudflare's own tables.
When should I use a decision model instead of an LLM?
Whenever the decision has a fixed set of answers and speed matters: ticket routing, content moderation, agent validation, next-action selection. If you need open-ended reasoning, prose, or tool calls, use a chat model — decision models do not generate text at all.
How do I try Clef?
Two paths: run it hosted on Workers AI (@cf/cloudflare/clef or @cf/cloudflare/clef-flash) with the Jev-compatible API, or download the weights from Hugging Face (Cloudflare/clef, Cloudflare/clef-flash) and serve them yourself — llama.cpp already supports decision models through its /v1/systemone endpoint.
A repeatable five-step workflow for turning any product into a finished ad poster with AI: lock the brief, choose the concept, structure the prompt, iterate in layers, and finalize type and composition.
Posters live or die on readable text — and most image models still mangle it. Here is how ChatGPT Images, Ideogram, and Recraft compare for poster work, which to pick for each job, and how to prompt text so it actually renders correctly.
The hardest part of AI storytelling is keeping one character looking like the same person across images. These five techniques — reference locking, description blocks, seed control, style anchors, and edit-based workflows — will keep your character recognizable from panel to panel.