Cloudflare released two open-weight decision models on 1 October 2026: Clef, a 27B model post-trained from Qwen3.8-27B, and Clef-Flash, a 9B model built on Qwen3.5-9B. Both are Apache 2.0 on Hugging Face and hosted on Workers AI.
A decision model doesn't write text. You send it a state (a support ticket, a JSON record, an image) and a list of typed questions, and it returns a probability for every allowed answer in one forward pass. That changes the bill. Workers AI lists Clef at $0.24 and Clef-Flash at $0.09 per million input tokens, and there is no output price at all. This post works out what that means for a ticket router running 100,000 decisions a month, where Clef-Flash falls short of the big model, and what you need to run either one yourself.
Table of Contents
- What a decision model returns
- The price sheet has no output column
- 100,000 ticket decisions, five models
- Clef or Clef-Flash: where the 9B falls short
- A router you can copy
- Running Clef on your own GPU
- Three catches the launch post skips
- FAQ
- Conclusion
- Sources
What a decision model returns
A normal LLM classifier works like this: you write a prompt asking for JSON, the model generates tokens one at a time, and you parse the result and hope the label is one you allowed. Clef skips the generation step. The Qwen backbone does a single prefill pass over your input, then a small joint schema head (4 transformer layers, 1,024 wide) reads the final hidden states and scores every allowed option of every question at once.
There are three question types:
| Type | What you send | What comes back |
|---|---|---|
noul | a yes/no question | the probability of yes |
choice | named options with descriptions | the top option, a probability for each, and a confidence |
score | an ordered list of levels | a probability-weighted score (it can land between levels), per-level probabilities, a confidence |
One request can carry 1 to 64 questions, and on Workers AI up to 4 images. Because the model can only score options you defined, it can't return a label outside your schema, and there is nothing to parse.
The price sheet has no output column
Workers AI bills in neurons at $0.011 per 1,000, and the pricing page converts each model to a per-token rate. For the two Clef models it lists input only:
| Model | Input per 1M tokens | Output per 1M tokens | Neurons per 1M input |
|---|---|---|---|
@cf/cloudflare/clef | $0.24 | none listed | 21,818 |
@cf/cloudflare/clef-flash | $0.09 | none listed | 8,182 |
@cf/qwen/qwen3.8-27b | $0.45 | $3.20 | 40,909 |
@cf/openai/gpt-oss-120b | $0.35 | $0.75 | 31,818 |
Clef's input rate is 53% of what the same platform charges for its own Qwen3.8-27B backbone. The bigger saving is the missing output line. An LLM classifier pays for every token of its JSON answer, and for every reasoning token before it if thinking is on. Clef's response schema still includes a usage.output_tokens field, but the pricing page has no rate to multiply it by.
Every account also gets 10,000 free neurons a day. At the rates above that is about 458,000 Clef input tokens or 1.22 million Clef-Flash input tokens per day, enough to prototype on real traffic before you pay anything.
100,000 ticket decisions, five models
Take a support inbox. Each decision sends about 1,200 input tokens (the ticket, some customer context and the question schema) and asks three questions: is it urgent, which team owns it, and how severe is it. An LLM doing the same job returns a small JSON object of about 60 tokens. Here is 100,000 decisions a month at Workers AI list prices:
| Model | Input cost | Output cost | Total per 100K | Per 1,000 decisions |
|---|---|---|---|---|
| Clef-Flash | $10.80 | $0.00 | $10.80 | $0.108 |
| Clef | $28.80 | $0.00 | $28.80 | $0.288 |
| gpt-oss-120b (60 tokens out) | $42.00 | $4.50 | $46.50 | $0.465 |
| Qwen3.8-27B (60 tokens out) | $54.00 | $19.20 | $73.20 | $0.732 |
| Qwen3.8-27B (600 tokens with thinking) | $54.00 | $192.00 | $246.00 | $2.460 |
Clef costs 2.5 times less than its own backbone run as a JSON classifier, and Clef-Flash costs 2.7 times less than Clef. The last row is the one people forget: turn on reasoning and let the model think for 600 tokens per ticket, and output becomes 78% of the bill. A decision model has no thinking budget to blow.
These figures assume the ticket is the whole input. If you attach long account histories, the input line grows for every model the same way, so the ranking holds. Plug your own token counts into the API Cost Calculator; both Clef models are in its model list now.
Clef or Clef-Flash: where the 9B falls short
Cloudflare's own Decision Index run puts Clef-Flash level with Clef on a lot of tasks, and ahead on several. BFCL function-call accuracy is 98.8 against 98.5. On a home-appliance tool simulator Flash scores 97.7 against 83.0, and its median latency is 38.8 ms against 209.3 ms.
The gaps open on tasks where the right answer is "none of the above" or "this is wrong":
| Benchmark (higher is better) | Clef | Clef-Flash | Gap |
|---|---|---|---|
| CLINC150+OOS intent detection, macro-F1 | 97.4 | 66.8 | 30.6 points |
| RAGTruth hallucination detection, F1 | 79.4 | 35.6 | 43.8 points |
| GSM8K | 80.8 | 67.3 | 13.5 points |
| BANKING77 intent, macro-F1 | 94.2 | 90.9 | 3.3 points |
CLINC150+OOS includes out-of-scope queries that match none of the intents, and Flash loses 30.6 points there. If your router gets messages that belong to no team, or you want a model to flag a RAG answer that isn't supported by its sources, use the 27B. For a closed set of labels where every input fits one of them, Flash at $0.09 is the better buy.
Neither model is a reasoner. On GPQA Diamond Clef scores 48.0 against 78.3 for Typesafe's Jev, and on BBH 73.7 against 92.9. Keep the hard thinking in an LLM and use Clef for the routing step in front of it.
All of these numbers come from Cloudflare's internal run of the Decision Index 0.2.1 suite, published alongside the models, so treat them as a vendor benchmark until someone else reproduces them.
A router you can copy
This is a Worker that sends every incoming ticket to Clef-Flash and only escalates to a human when the model isn't confident. The request shape is from Cloudflare's model page; the thresholds are a starting point to tune on your own data.
export interface Env {
AI: Ai;
}
export default {
async fetch(request, env): Promise<Response> {
const ticket = await request.text();
const result = await env.AI.run('@cf/cloudflare/clef-flash', {
model: 'clef-flash',
state: ticket,
questions: {
urgent: { type: 'noul', instructions: 'Is this support request urgent?' },
team: {
type: 'choice',
instructions: 'Which team should handle this request?',
criteria: {
billing: 'Payments, invoices, and refunds',
technical: 'Outages, errors, and configuration',
sales: 'Plans and upgrades',
},
},
severity: {
type: 'score',
instructions: 'How severe is the customer impact?',
criteria: ['No impact', 'Minor', 'Major', 'Critical'],
},
},
});
const { urgent, team, severity } = result.answers;
// Low confidence on the team means the ticket may fit no queue: send it to a person.
const needsHuman = team.confidence < 0.7;
const page = urgent.noul > 0.8 && severity.score >= 2;
return Response.json({ team: team.choice, page, needsHuman, usage: result.usage });
},
} satisfies ExportedHandler<Env>;Two habits help. Log the full probabilities object for a week before you pick thresholds, since a fixed 0.7 rarely survives contact with real traffic. And write option descriptions the way you would brief a new hire, because the head scores your criteria text against the state. "Outages, errors, and configuration" routes better than "tech".
Running Clef on your own GPU
The weights are plain BF16 safetensors plus a separate joint_head.safetensors. From the model's config.json:
| Clef | Clef-Flash | |
|---|---|---|
| Base model | Qwen3.8-27B | Qwen3.5-9B |
| Parameters (backbone with vision) | 27.36B | 9.41B |
| BF16 weights | 54.7 GB (51.0 GiB) | 18.8 GB (17.5 GiB) |
| Layers | 64, of which 16 full attention | 32, of which 8 full attention |
| KV heads, head dim | 4, 256 | 4, 256 |
| KV cache per token (BF16) | 64 KiB | 32 KiB |
| KV cache at the 16,384-token default | 1.0 GiB | 0.5 GiB |
Both use Qwen's 3:1 hybrid stack, where three linear-attention layers sit between every full-attention layer and only the full-attention ones keep a growing KV cache. That keeps the cache small. And because Clef only prefills, the cache never grows past your input; there are no generated tokens to add.
So Clef-Flash at BF16 fits a 24 GB card with room for the cache and runtime. The 27B needs an 80 GB-class GPU at BF16; Cloudflare tested it on a single H200 with torch 2.11 and transformers 5.10.2. Both models are presets in the AI VRAM Calculator, with only the full-attention layers counted for the cache, so you can try other context lengths.
Local tools are a weaker fit. The scoring head is custom PyTorch code (joint_schema_model.py) that the model card loads by adding the download folder to sys.path, so you can't drop Clef into Ollama or LM Studio the way you would a chat model. If you already run a local agent stack such as the one in our Hermes Agent setup guide, plan on a small Python service next to it.
Three catches the launch post skips
The context window depends on where you run it. Workers AI lists 65,536 tokens, the announcement says 64k, and the backbone config allows 262,144. But the open-source encode_record helper defaults to max_length 16,384 tokens. Run the weights yourself without changing that and long states get cut at 16K. Workers AI also truncates long text state to fit, so check usage.input_tokens against what you sent.
You still have to check calibration yourself. Cloudflare trained with a Brier loss and its own RL method to calibrate the probabilities, and Clef's ForecastBench Brier score of 13.9 beats Jev's 17.4. A 0.9 from Clef on your data only means 90% once a labelled sample of your own tickets agrees. Those 381 free Clef decisions a day are enough to build that sample.
The benchmarks are first-party. Every score in this post comes from Cloudflare's run of a suite designed around Jev, Typesafe's decision model. Clef is API-compatible with Jev and SystemOne, so the cheapest test is to send the same 1,000 real tickets to both and compare.
FAQ
How much does Cloudflare Clef cost? On Workers AI, Clef is $0.24 and Clef-Flash $0.09 per million input tokens, with no output rate listed. A 1,200-token decision costs about $0.000288 on Clef and $0.000108 on Clef-Flash.
Is Clef open source? Yes. Both models are on Hugging Face under Apache 2.0, the same license as their Qwen base models.
Can Clef generate text or call tools? No. It only returns probabilities for the options you define. Use it to decide which tool, team or branch to take, and hand any writing to an LLM.
Does Clef read images? Yes. It keeps Qwen's vision encoder, so a state can include images or video frames. Workers AI accepts up to 4 images per request.
Can Clef-Flash run on a 24 GB GPU? Yes at BF16. The weights are 17.5 GiB and the KV cache at 16K tokens is 0.5 GiB. The 27B Clef needs about 51 GiB for weights alone.
Conclusion
Clef turns classification into an input-only bill. For a ticket router at 100,000 decisions a month, Clef-Flash costs $10.80 and Clef $28.80, against $73.20 for the same Qwen backbone answering in JSON and $246 if you let it think. Use Clef-Flash when every input fits a known label, move to Clef when out-of-scope inputs or hallucination checks matter, and keep an LLM for the steps that need real reasoning. Before you switch, calibrate the probabilities on your own labelled tickets, and if you self-host, raise max_length past the 16K default.
Sources
- Cloudflare blog, "Introducing Clef: our open-source decision models, and new RL fine-tuning platform" (1 October 2026): decision-model concept, 64k context, training method, Jev comparison, Typesafe workflow evals, latency.
- Cloudflare Docs, Workers AI pricing (last updated 1 October 2026): Clef, Clef-Flash, Qwen3.8-27B and gpt-oss-120b rates, neuron conversions, the $0.011 per 1,000 neurons price and 10,000 free neurons per day.
- Cloudflare Docs,
@cf/cloudflare/clefand@cf/cloudflare/clef-flashmodel pages: 65,536-token context, request parameters, 64-question and 4-image limits, output schema. - Hugging Face,
Cloudflare/clefandCloudflare/clef-flashmodel cards andconfig.json: base models, joint schema head, Decision Index 0.2.1 results,max_lengthdefault, parameter counts and layer layout. - Decision costs and VRAM figures are calculated by ToolMintX from those rates and configs, using 1,200 input tokens per decision and 60 or 600 output tokens for the LLM rows.
