The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Dispatches · Cloudflare · Perplexity · Ollama · Qwen

Cloudflare and Perplexity release open 27B decision models built on Qwen

Cloudflare's Clef and Perplexity's pplx-decider both turn a 27-billion-parameter Qwen model into a classifier for routing tickets and labelling bug reports, and Ollama now lists Clef to pull in one command.

Two companies have put out open models built for decisions rather than conversation. Cloudflare's Clef is listed on Ollama, which says it can classify an image, label a bug report or route a support ticket to the right team. Ollama's post names two sizes: Clef at 27 billion parameters and Clef Flash at 9 billion. Both pull with a single command.

According to the Ollama listing, Clef takes a state and a schema of typed questions and returns decisions. It accepts text, JSON or images, including screenshots, receipts and forms, in a single pass. The listing gives a 256,000-token context window and an 18GB download for the 27B version. It requires Ollama 0.35.1 or later and is released under Apache 2.0.

Cloudflare, per the Ollama page, fine-tuned Clef from Qwen 3.8-27B and says it tops what it calls the Decision Index, beating rival models including one named Jev. That is the maker's own claim. This report has not found an independent run of the index, and Ollama's page does not say who maintains it.

Perplexity has a sibling. Its model card for pplx-decider-v1-27b describes a classifier also fine-tuned from Qwen3.8-27B, under Apache 2.0, taking text and images. Perplexity says it reaches 85.71 per cent accuracy across 11 benchmarks, against 74.76 per cent for the base model, with 90.60 per cent on TabFact. Hugging Face lists it as the 73rd most trending model, with 64 likes.

Neither company has published head-to-head numbers against the other, and no outside tester has reported on either model yet. The card says the Perplexity model needs about 49 GiB of GPU memory in Python, a heavier lift than Clef's quantised download. A llama.cpp endpoint for decision models was reported earlier this week, which suggests tooling is moving ahead of the benchmarks.

Sources 3 sources

  1. Source Ollama
  2. Source Perplexity on Hugging Face
  3. Source Ollama on X