At 01:30 UTC on October 2, Cloudflare's new Clef model had 336 Hugging Face likes and just 18 downloads in MrKeyoor's source snapshot. It had also surfaced on Product Hunt. Developers were interested in the idea before many had run the weights. Clef takes a state, a fixed set of questions and their allowed answers, then scores every option without generating a sentence or parsing JSON afterward.
Cloudflare released two versions on October 1: the 27-billion-parameter Clef and the smaller Clef-flash. Both are available through Workers AI, and their weights are published under Apache 2.0. The Clef model card says the larger model accepts text, JSON, images and video, then returns probabilities for typed questions in one forward pass. For an agent that only needs to choose a queue, score severity or decide whether to call a tool, the model is designed to stop before prose begins.
A model call that ends before writing
Most structured LLM calls still ask a language model to write something. Even with constrained decoding, the model emits tokens for a label or JSON object, and application code checks the result. Clef uses a Qwen3.8-27B backbone for a prefill pass, then scores the permitted schema choices in parallel. Cloudflare's technical account describes the decision stage as non-autoregressive: there is no hidden explanation or intermediate answer produced token by token.
The documented request format makes that boundary visible. A developer supplies a state plus questions of three types: noul for true or false, choice for named alternatives, and score for an ordered scale. The response contains probabilities for the supplied options. A checkout incident, for example, could be sent with questions about urgency, owning team and severity. Clef cannot create a fourth team that the schema omitted. It can still assign a high probability to the wrong one.
Questions are scored together rather than as isolated classifier calls. According to the published model files and architecture notes, a small transformer head reads the backbone's final hidden states, routes evidence to each question and lets fields attend to one another before assigning one logit to every option. That matters when answers are related. A ticket classed as a full outage should influence its severity score, for instance, without requiring one model response to be fed into another.
Cloudflare froze the Qwen backbone and trained the routing head alongside rank-256 low-rank adapters. Its post-training combined label-smoothed cross-entropy with a Brier loss intended to improve probability calibration. The company also describes a reinforcement-learning objective that gives partial credit to nearby ordered choices and rewards an entirely correct record. Those are Cloudflare's training claims. The release does not include an independent audit of the training data or calibration.
The benchmark table has real losses
Cloudflare reports strong numbers on several fixed-choice tasks. In its Decision Index run, Clef scored 98.5 percent exact accuracy on BFCL, 91.9 percent on API-Bank and a 94.2 macro-F1 on BANKING77. Clef-flash reached 98.8 percent, 93.1 percent and 90.9 percent on the same tasks. The full table on the model card identifies these as results from Cloudflare's internal run, so they have not been independently reproduced.
Latency is the sharper result. The published median was 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash and 524.1 milliseconds for Jev. Laya recorded 5.8 milliseconds, far faster than every other entry, while scoring much lower on BFCL, API-Bank and BANKING77. The table describes a speed and quality trade, rather than a single winner. Hardware and serving configuration need to travel with any latency claim before a team treats those numbers as a capacity plan.
Clef also loses some comparisons in the same published results. Jev scored 81.0 percent on When2Call against Clef's 72.4 percent, and 82.7 percent on MMLU-Pro against 65.9 percent. In Cloudflare's four workflow evaluations, Clef led invoice processing and security-incident exact actions, Clef-flash narrowly led customer service, and Jev led agent-trace observability. A routing model can look excellent on one taxonomy and mediocre on another. That variation is more informative than the claim that one model tops an aggregate leaderboard.
Cloudflare tested a more concrete internal path with its threat-intelligence team. In the company's reported Browser Run test, Clef fetched, rendered and classified a website in 2.2 seconds, while its gpt-oss-120b workflow took 4.7 seconds and returned only two classifications. This is a company-run example that mixes browser time with inference time, so it should not be read as a general two-times model-speed result. It does show the intended job: produce several bounded signals about one state and pass them directly to ordinary software.
The large model still needs large hardware
Skipping generation does not make a 27B model small. Cloudflare tested the downloadable Clef weights with PyTorch 2.11 and Transformers 5.10.2 on a single Nvidia H200, according to the usage section. The release uses BF16 safetensors and includes the Qwen vision encoder, the joint schema head, tokenizer and media processor. The page also links community quantizations for llama.cpp, Ollama and LM Studio, but Cloudflare does not publish equivalent quality and latency results for those variants.
Clef-flash is the practical counterpoint. It uses a frozen Qwen3.5-9B backbone and the same basic typed-decision interface. Its 38.8 millisecond median in Cloudflare's comparison was about one-fifth of Clef's, and it beat the larger model on BFCL, API-Bank, the home-appliance simulator and several other tasks. Parameter count did not settle the choice. Developers need to test the exact schema and error cost their application carries.
The open release is still useful even for teams that never serve 27B parameters. The Apache-2.0 weights and Python implementation expose the record encoding, batching, routing head and probability calculation. Researchers can inspect the mechanism, rerun the evaluation and test calibration on another dataset. Workers AI offers the hosted path, while the downloadable code keeps the core decision interface from being available only as an API claim.
Typed output removes one failure class
A bounded answer is easier to connect to code than generated prose. Under the Clef API contract, there is no malformed JSON to repair, no invented label to map back to a queue and no variable number of answers. The application can compare a probability with a threshold, send uncertain cases to review and record the complete distribution. Clef's interface is especially suited to repeated decisions whose permissible outcomes are known before inference.
The fixed schema moves responsibility into the application. If the real answer is absent, Clef must distribute probability among wrong choices. If two categories are poorly defined, a well-calibrated probability cannot repair the taxonomy. A team also needs held-out examples from its own traffic to set thresholds. A confidence value produced by the model is evidence for a policy decision, not permission to remove the fallback path.
Multimodal input widens the possible uses and the testing burden. The model card's example pairs a receipt image with a question about whether its total is legible, and text-only and multimodal records can share a batch. Video arrives as frame arrays. A system routing visual reports or moderating uploaded media must measure what resizing, frame selection and degraded images do to its probabilities. The release's headline benchmark table cannot answer those product-specific questions.
Open weights sit beside a hosted training business
Clef is also the first Cloudflare-trained model from the Workers AI team and the starting point for a customer fine-tuning service. Cloudflare says forward-deployed engineers will initially help customers adapt the model. The company plans a later self-serve system built from AI Gateway datasets, Workers AI rollouts, Containers for scoring and replay, a new Trainer component, and deployment back to Workers AI. Several of those pieces are described as work in progress in the launch post.
That commercial path clarifies Cloudflare's bet. The model supplies a fixed decision surface. The hosted plan collects workload examples, tunes the weights and serves the result close to application code. Cloudflare is placing decision models in the hot path between an agent gathering context and an LLM or tool taking action. Whether that split improves a real system depends on the errors at the handoff. Median classifier latency covers only one part of the path.
The 336 likes and 18 downloads in the initial Hugging Face snapshot measure curiosity. Adoption is still unproven. The next useful evidence will be independently reproduced latency on named hardware, calibration curves on unfamiliar datasets and failure rates from complete workflows where one wrong route changes the next action. Clef's most testable promise is narrow: when software already knows the possible answers, the model can score them without writing around them first. Developers now need to measure how often the right answer survives that shortcut.