Skip to content

AI·News & analysis

Cloudflare releases Clef, open-weight AI models that make yes-or-no decisions fast

Clef and Clef-flash answer bounded questions for AI agents in milliseconds instead of writing text. Cloudflare hosts them and also gives away the weights, a direct challenge to TypeSafe's Jev.

By Dan Kost aka Poseidan8 min read
Rows of colorful glowing lava lamps on white shelves at a Cloudflare office
Photo: HaeB / Wikimedia Commons, CC BY-SA 4.0

Tide

Ripple

Sci-fi

5/10

Reality

In your hands

An AI that skips the speech and just answers yes or no.How we rate

The Squeeze

Cloudflare released Clef and Clef-flash, open-weight decision models that give AI agents quick structured answers instead of free-form text.

They are hosted on Workers AI and downloadable from Hugging Face under Apache 2.0. Cloudflare says they beat TypeSafe's Jev on speed and most of its benchmarks, though those results are self-reported. OpenAI and Amazon launched similar models days earlier.

What to know

  1. Cloudflare released Clef and Clef-flash, its first models trained by the Workers AI team, built to answer yes/no, multiple-choice and ranking questions for AI agents.
  2. Both are hosted on Workers AI and open-weight under Apache 2.0 on Hugging Face, and they are API-compatible with TypeSafe's Jev.
  3. Cloudflare says Clef is 2.5 times faster than Jev at the median and Clef-flash 13 times faster, across 43 benchmark runs; these are self-reported scores.
  4. Hosted Clef costs $0.24 per million tokens, nearly six times Jev's price, The Register reports, but you can run it yourself on a GPU with enough memory.

Cloudflare released Clef and Clef-flash, the first AI models trained by its own Workers AI team. The fun question: why would an AI agent want a model that refuses to write sentences?

First, what's a decision model?

Most AI chatbots answer by writing text, one word at a time. That's great for essays. It's slow and wasteful when an agent just needs a yes or a no.

A decision model skips the writing. Cloudflare explains it this way: Clef reads an input and a set of typed questions, then returns a probability for every allowed answer. The agent gets a structured decision it can act on right away, like "route this ticket" or "block this request."

Think of it like a multiple-choice exam versus an essay exam. The decision model just fills in the bubble.

Clef answers three kinds of questions, according to Cloudflare:

  • Yes/no: the probability the answer is yes.
  • Choice: pick one option from a list you define.
  • Score: rate something against a scale you set.

You can ask up to 64 questions in a single request.

Why is everyone suddenly building these?

This category took off fast. The Register says Clef arrived two weeks after TypeSafe's Jev model "took the AI world by storm."

Startup Fortune explains the appeal with one example: a hackathon demo reportedly used Jev to monitor an AI agent's actions for $2.94, compared with $372 running the same check through a big frontier model.

Clef is also not alone this week. Startup Fortune reports that OpenAI introduced a Decisions API at its Dev Day on September 30, and Amazon released its own small decision model, Strands Decider 2B, on October 1.

Here's how the three compare, based on Startup Fortune's reporting:

Model Maker Weights
Decisions API (on Luna) OpenAI Hosted only, inside OpenAI's stack
Strands Decider 2B Amazon Open, one small 2B size
Clef and Clef-flash Cloudflare Open, two sizes (27B and 9B)

Of the three, Cloudflare is the only one giving away the model weights at two sizes, Startup Fortune notes.

Startup Fortune adds some detail on Amazon's entry. Strands Decider 2B came out of Amazon's Strands Labs and is built on a Qwen3.5 base. It reportedly started as an internal prototype from AWS distinguished engineer Marc Brooker, and scores around 72% on the public JevBench v19 test set.

Its takeaway for developers: the real cost isn't only the price per answer, it's lock-in. A team that builds on a closed API depends on one company's uptime and pricing. Open weights let you run the model wherever you want.

How do Clef and Clef-flash differ?

  • Clef: a 27 billion parameter model aimed at the most accurate answers.
  • Clef-flash: a 9 billion parameter model built to answer in a few tens of milliseconds.

The Register reports both are built on frozen, post-trained versions of Alibaba's Qwen models. Instead of generating words, they read the question once and score the possible answers directly. Startup Fortune calls this a "non-autoregressive, prefill-only" approach.

That's a contrast with Jev, whose design TypeSafe has kept secret, The Register notes.

Both Clef models support a 64k token context window. Jev also handles up to 64k tokens per request, The Register adds, but its state plus its longest single question is capped at 32k. Clef can also look at images: Cloudflare says it has a vision encoder and accepts up to four images per request. The Register adds that it handles video too, something Jev doesn't.

Is it really faster than Jev?

By Cloudflare's own numbers, yes. Across 43 benchmark runs, Cloudflare says Clef was 2.5 times faster than Jev at the median, and Clef-flash 13 times faster.

On accuracy, Cloudflare says a Clef model scored highest on 7 of 10 decision benchmarks. On TypeSafe's own workflow tests, Clef beat Jev in 3 of 4 areas: invoice processing, customer service and security incidents.

The catch: The Register points out that Cloudflare self-reported those scores. They hadn't yet been reproduced for the official Decision Index ranking. The one area Clef lost, The Register adds, was agent trace observability. The Register also notes a trade-off in Cloudflare's own ranking: Clef is a little slower than some other open decision models but more accurate, while Clef-flash matches most of them on accuracy and is far faster.

One real-world test from Cloudflare: paired with its Browser Run tool, Clef fetched, rendered and classified a website in 2.2 seconds, versus 4.7 seconds for the gpt-oss-120b model in the same workflow.

What does it cost to use?

You have two options.

Hosted: Clef runs on Cloudflare's Workers AI. Cloudflare says that adds speed, because requests run on GPUs across its network, close to users. The Register reports the price: $0.24 per million tokens, nearly six times Jev's $0.042.

Self-hosted: the weights are on Hugging Face under the Apache 2.0 license, so you can run Clef on your own hardware. Cloudflare's Michelle Chen told The Register what you'll need:

Model GPU memory needed
Clef-flash At least 41 GB
Clef 85 GB

Chen said that assumes one request at a time and a 64k context window. She also confirmed to The Register that the training datasets aren't public, even though the announcement called the models open source.

Switching should be easy for Jev users. Cloudflare says Clef follows the same API, so you change the endpoint and model name.

Developers can call Clef through the Workers AI binding in code or a REST API, and route it through Cloudflare's AI Gateway. Cloudflare's suggested pattern is to put Clef directly in an agent's request path for the quick decision, then hand off to a larger language model on Workers AI when the agent needs to actually do something.

What can you build with it?

Cloudflare lists several uses:

  • Support triage: decide if a ticket is urgent and which team owns it, then route it without a human.
  • Trust and safety: score user posts against your own policy rules.
  • Agent guardrails: let an agent ask "should I take this action?" in tens of milliseconds before calling a tool.
  • Threat intelligence: classify websites by category.
  • Visual checks: classify images, not just text.

In real life A customer writes "checkout has been failing for every customer for the last hour." Clef instantly marks it urgent and sends it to the technical team, before any human reads it.

Cloudflare is also launching a reinforcement learning fine-tuning service so customers can tune Clef for their own workloads. It's asking interested teams to sign up as design partners. Cloudflare had a busy week: it also made its Basin data platform generally available.

What it means for you

  • If you build AI agents: you can try Clef on Workers AI today, or download it and run it yourself.
  • If you already use Jev: Clef is a drop-in alternative, but compare prices, since hosted Clef costs more per token.
  • If you care about lock-in: open weights mean you can keep running the model even if you leave Cloudflare's hosting.
  • If you're just curious: these models are why more apps may soon make small decisions for you, quickly and quietly.

The bottom line

Cloudflare joined the race for fast decision models with Clef and Clef-flash, and it's the only big player so far giving away the weights at two sizes. Its speed and accuracy claims look strong, but they're self-reported until independent tests catch up.

Key facts

Models
Clef (27B) and Clef-flash (9B)
License
Apache 2.0, weights on Hugging Face
Speed claim
2.5x (Clef) and 13x (Clef-flash) faster than Jev at the median
Hosted price
$0.24 per million tokens (The Register)
Context window
64k tokens

Got questions?

Quick answers, plain words

What is Cloudflare Clef?

Clef is a decision model from Cloudflare. Instead of writing text, it reads some input and a set of typed questions, then returns a probability for each allowed answer, so an AI agent can act right away.

What is a decision model?

It's an AI model built to answer bounded questions, like yes or no, pick one option, or rate on a scale, rather than write open-ended text. TypeSafe's Jev started this category.

Is Clef open source?

The weights are open under the Apache 2.0 license on Hugging Face. Cloudflare's Michelle Chen told The Register the training datasets are not public.

What's the difference between Clef and Clef-flash?

Clef is a 27 billion parameter model aimed at accuracy. Clef-flash has 9 billion parameters and is built to answer in a few tens of milliseconds.

How much does Clef cost?

Hosted on Workers AI, Clef costs $0.24 per million tokens, according to The Register, nearly six times Jev's $0.042. Running it yourself from Hugging Face avoids the token price.

What hardware do I need to run Clef locally?

Cloudflare's Michelle Chen told The Register that Clef-flash needs a GPU with at least 41 GB of memory and Clef needs 85 GB, assuming one request at a time and a 64k context window.

Can Clef replace Jev?

Cloudflare says Clef follows the same API, so a Jev integration can switch by changing the endpoint and model name.

Can Clef handle images?

Yes. Clef has a vision encoder and can take up to four images with a request, Cloudflare says. The Register adds that it can handle video too.

Are Cloudflare's benchmark results independent?

No. The Register notes Cloudflare self-reported its scores, and they had not yet been reproduced on the official Decision Index ranking.

SourcesCloudflare
Topics and tagsAI agents, cloudflare, ai agents, open source

The daily newsletter

Tech news you'll actually get.

One short email a day. Five minutes. Plain words. The daily email is launching soon. Join the early list.

Free. Early list: we'll email you when the first issue goes out.

More in brief