Editorial

Cloudflare Clef: Decision Models, No Output Bill

Cloudflare Clef bills input only: $0.24/M, or $0.09/M for Clef-Flash. 100K ticket decisions cost $10.80 vs $73.20 for an LLM.

JJyoti Ranjan SwainUpdated
Cloudflare Clef and Clef-Flash input-only pricing: $0.24 and $0.09 per million tokens, with no output rate

Cloudflare released two open-weight decision models on 1 October 2026: Clef, a 27B model post-trained from Qwen3.8-27B, and Clef-Flash, a 9B model built on Qwen3.5-9B. Both are Apache 2.0 on Hugging Face and hosted on Workers AI.

A decision model doesn't write text. You send it a state (a support ticket, a JSON record, an image) and a list of typed questions, and it returns a probability for every allowed answer in one forward pass. That changes the bill. Workers AI lists Clef at $0.24 and Clef-Flash at $0.09 per million input tokens, and there is no output price at all. This post works out what that means for a ticket router running 100,000 decisions a month, where Clef-Flash falls short of the big model, and what you need to run either one yourself.

Table of Contents

What a decision model returns

A normal LLM classifier works like this: you write a prompt asking for JSON, the model generates tokens one at a time, and you parse the result and hope the label is one you allowed. Clef skips the generation step. The Qwen backbone does a single prefill pass over your input, then a small joint schema head (4 transformer layers, 1,024 wide) reads the final hidden states and scores every allowed option of every question at once.

Clef request flow: state and typed questions pass through one Qwen prefill and a schema head that outputs probabilities

There are three question types:

TypeWhat you sendWhat comes back
noula yes/no questionthe probability of yes
choicenamed options with descriptionsthe top option, a probability for each, and a confidence
scorean ordered list of levelsa probability-weighted score (it can land between levels), per-level probabilities, a confidence

One request can carry 1 to 64 questions, and on Workers AI up to 4 images. Because the model can only score options you defined, it can't return a label outside your schema, and there is nothing to parse.

The price sheet has no output column

Workers AI bills in neurons at $0.011 per 1,000, and the pricing page converts each model to a per-token rate. For the two Clef models it lists input only:

ModelInput per 1M tokensOutput per 1M tokensNeurons per 1M input
@cf/cloudflare/clef$0.24none listed21,818
@cf/cloudflare/clef-flash$0.09none listed8,182
@cf/qwen/qwen3.8-27b$0.45$3.2040,909
@cf/openai/gpt-oss-120b$0.35$0.7531,818

Clef's input rate is 53% of what the same platform charges for its own Qwen3.8-27B backbone. The bigger saving is the missing output line. An LLM classifier pays for every token of its JSON answer, and for every reasoning token before it if thinking is on. Clef's response schema still includes a usage.output_tokens field, but the pricing page has no rate to multiply it by.

Every account also gets 10,000 free neurons a day. At the rates above that is about 458,000 Clef input tokens or 1.22 million Clef-Flash input tokens per day, enough to prototype on real traffic before you pay anything.

100,000 ticket decisions, five models

Take a support inbox. Each decision sends about 1,200 input tokens (the ticket, some customer context and the question schema) and asks three questions: is it urgent, which team owns it, and how severe is it. An LLM doing the same job returns a small JSON object of about 60 tokens. Here is 100,000 decisions a month at Workers AI list prices:

ModelInput costOutput costTotal per 100KPer 1,000 decisions
Clef-Flash$10.80$0.00$10.80$0.108
Clef$28.80$0.00$28.80$0.288
gpt-oss-120b (60 tokens out)$42.00$4.50$46.50$0.465
Qwen3.8-27B (60 tokens out)$54.00$19.20$73.20$0.732
Qwen3.8-27B (600 tokens with thinking)$54.00$192.00$246.00$2.460

Monthly cost of 100,000 support-ticket decisions: Clef-Flash $10.80, Clef $28.80, gpt-oss-120b $46.50, Qwen3.8-27B $73.20

Clef costs 2.5 times less than its own backbone run as a JSON classifier, and Clef-Flash costs 2.7 times less than Clef. The last row is the one people forget: turn on reasoning and let the model think for 600 tokens per ticket, and output becomes 78% of the bill. A decision model has no thinking budget to blow.

These figures assume the ticket is the whole input. If you attach long account histories, the input line grows for every model the same way, so the ranking holds. Plug your own token counts into the API Cost Calculator; both Clef models are in its model list now.

Clef or Clef-Flash: where the 9B falls short

Cloudflare's own Decision Index run puts Clef-Flash level with Clef on a lot of tasks, and ahead on several. BFCL function-call accuracy is 98.8 against 98.5. On a home-appliance tool simulator Flash scores 97.7 against 83.0, and its median latency is 38.8 ms against 209.3 ms.

The gaps open on tasks where the right answer is "none of the above" or "this is wrong":

Benchmark (higher is better)ClefClef-FlashGap
CLINC150+OOS intent detection, macro-F197.466.830.6 points
RAGTruth hallucination detection, F179.435.643.8 points
GSM8K80.867.313.5 points
BANKING77 intent, macro-F194.290.93.3 points

CLINC150+OOS includes out-of-scope queries that match none of the intents, and Flash loses 30.6 points there. If your router gets messages that belong to no team, or you want a model to flag a RAG answer that isn't supported by its sources, use the 27B. For a closed set of labels where every input fits one of them, Flash at $0.09 is the better buy.

Neither model is a reasoner. On GPQA Diamond Clef scores 48.0 against 78.3 for Typesafe's Jev, and on BBH 73.7 against 92.9. Keep the hard thinking in an LLM and use Clef for the routing step in front of it.

All of these numbers come from Cloudflare's internal run of the Decision Index 0.2.1 suite, published alongside the models, so treat them as a vendor benchmark until someone else reproduces them.

A router you can copy

This is a Worker that sends every incoming ticket to Clef-Flash and only escalates to a human when the model isn't confident. The request shape is from Cloudflare's model page; the thresholds are a starting point to tune on your own data.

ts
export interface Env {
  AI: Ai;
}

export default {
  async fetch(request, env): Promise<Response> {
    const ticket = await request.text();

    const result = await env.AI.run('@cf/cloudflare/clef-flash', {
      model: 'clef-flash',
      state: ticket,
      questions: {
        urgent: { type: 'noul', instructions: 'Is this support request urgent?' },
        team: {
          type: 'choice',
          instructions: 'Which team should handle this request?',
          criteria: {
            billing: 'Payments, invoices, and refunds',
            technical: 'Outages, errors, and configuration',
            sales: 'Plans and upgrades',
          },
        },
        severity: {
          type: 'score',
          instructions: 'How severe is the customer impact?',
          criteria: ['No impact', 'Minor', 'Major', 'Critical'],
        },
      },
    });

    const { urgent, team, severity } = result.answers;
    // Low confidence on the team means the ticket may fit no queue: send it to a person.
    const needsHuman = team.confidence < 0.7;
    const page = urgent.noul > 0.8 && severity.score >= 2;

    return Response.json({ team: team.choice, page, needsHuman, usage: result.usage });
  },
} satisfies ExportedHandler<Env>;

Two habits help. Log the full probabilities object for a week before you pick thresholds, since a fixed 0.7 rarely survives contact with real traffic. And write option descriptions the way you would brief a new hire, because the head scores your criteria text against the state. "Outages, errors, and configuration" routes better than "tech".

Running Clef on your own GPU

The weights are plain BF16 safetensors plus a separate joint_head.safetensors. From the model's config.json:

ClefClef-Flash
Base modelQwen3.8-27BQwen3.5-9B
Parameters (backbone with vision)27.36B9.41B
BF16 weights54.7 GB (51.0 GiB)18.8 GB (17.5 GiB)
Layers64, of which 16 full attention32, of which 8 full attention
KV heads, head dim4, 2564, 256
KV cache per token (BF16)64 KiB32 KiB
KV cache at the 16,384-token default1.0 GiB0.5 GiB

Both use Qwen's 3:1 hybrid stack, where three linear-attention layers sit between every full-attention layer and only the full-attention ones keep a growing KV cache. That keeps the cache small. And because Clef only prefills, the cache never grows past your input; there are no generated tokens to add.

Clef and Clef-Flash VRAM at BF16: 51.0 GiB and 17.5 GiB of weights plus 1.0 and 0.5 GiB of KV cache at 16K tokens

So Clef-Flash at BF16 fits a 24 GB card with room for the cache and runtime. The 27B needs an 80 GB-class GPU at BF16; Cloudflare tested it on a single H200 with torch 2.11 and transformers 5.10.2. Both models are presets in the AI VRAM Calculator, with only the full-attention layers counted for the cache, so you can try other context lengths.

Local tools are a weaker fit. The scoring head is custom PyTorch code (joint_schema_model.py) that the model card loads by adding the download folder to sys.path, so you can't drop Clef into Ollama or LM Studio the way you would a chat model. If you already run a local agent stack such as the one in our Hermes Agent setup guide, plan on a small Python service next to it.

Three catches the launch post skips

The context window depends on where you run it. Workers AI lists 65,536 tokens, the announcement says 64k, and the backbone config allows 262,144. But the open-source encode_record helper defaults to max_length 16,384 tokens. Run the weights yourself without changing that and long states get cut at 16K. Workers AI also truncates long text state to fit, so check usage.input_tokens against what you sent.

You still have to check calibration yourself. Cloudflare trained with a Brier loss and its own RL method to calibrate the probabilities, and Clef's ForecastBench Brier score of 13.9 beats Jev's 17.4. A 0.9 from Clef on your data only means 90% once a labelled sample of your own tickets agrees. Those 381 free Clef decisions a day are enough to build that sample.

The benchmarks are first-party. Every score in this post comes from Cloudflare's run of a suite designed around Jev, Typesafe's decision model. Clef is API-compatible with Jev and SystemOne, so the cheapest test is to send the same 1,000 real tickets to both and compare.

FAQ

How much does Cloudflare Clef cost? On Workers AI, Clef is $0.24 and Clef-Flash $0.09 per million input tokens, with no output rate listed. A 1,200-token decision costs about $0.000288 on Clef and $0.000108 on Clef-Flash.

Is Clef open source? Yes. Both models are on Hugging Face under Apache 2.0, the same license as their Qwen base models.

Can Clef generate text or call tools? No. It only returns probabilities for the options you define. Use it to decide which tool, team or branch to take, and hand any writing to an LLM.

Does Clef read images? Yes. It keeps Qwen's vision encoder, so a state can include images or video frames. Workers AI accepts up to 4 images per request.

Can Clef-Flash run on a 24 GB GPU? Yes at BF16. The weights are 17.5 GiB and the KV cache at 16K tokens is 0.5 GiB. The 27B Clef needs about 51 GiB for weights alone.

Conclusion

Clef turns classification into an input-only bill. For a ticket router at 100,000 decisions a month, Clef-Flash costs $10.80 and Clef $28.80, against $73.20 for the same Qwen backbone answering in JSON and $246 if you let it think. Use Clef-Flash when every input fits a known label, move to Clef when out-of-scope inputs or hallucination checks matter, and keep an LLM for the steps that need real reasoning. Before you switch, calibrate the probabilities on your own labelled tickets, and if you self-host, raise max_length past the 16K default.

Sources

  • Cloudflare blog, "Introducing Clef: our open-source decision models, and new RL fine-tuning platform" (1 October 2026): decision-model concept, 64k context, training method, Jev comparison, Typesafe workflow evals, latency.
  • Cloudflare Docs, Workers AI pricing (last updated 1 October 2026): Clef, Clef-Flash, Qwen3.8-27B and gpt-oss-120b rates, neuron conversions, the $0.011 per 1,000 neurons price and 10,000 free neurons per day.
  • Cloudflare Docs, @cf/cloudflare/clef and @cf/cloudflare/clef-flash model pages: 65,536-token context, request parameters, 64-question and 4-image limits, output schema.
  • Hugging Face, Cloudflare/clef and Cloudflare/clef-flash model cards and config.json: base models, joint schema head, Decision Index 0.2.1 results, max_length default, parameter counts and layer layout.
  • Decision costs and VRAM figures are calculated by ToolMintX from those rates and configs, using 1,200 input tokens per decision and 60 or 600 output tokens for the LLM rows.

Tools In This Article

Browser-based, no sign-up. Try them while the topic is fresh.

More From ToolMintX

Other Blog Posts