llama.cpp adds an endpoint for decision models
Georgi Gerganov says the latest llama.cpp builds serve decision models through a new /v1/systemone endpoint, and ggml-org has posted Kev-4B, a 4-billion-parameter model under Apache 2.0.
Georgi Gerganov, who created llama.cpp, said on X on Friday that decision models are now available in the project's latest builds. In his post he describes a `/v1/systemone` endpoint for running this kind of inference "locally" and "privately", and says several open models are supported, with more to come. The post had about 450 likes when we read it.
https://x.com/ggerganov/status/2106029758350032937
Hugging Face chief executive Clément Delangue posted the same day that decision models now run on device in llama.cpp, with the command `llama serve -hf ggml-org/Kev-4B-GGUF`. The model card on Hugging Face lists Kev-4B as a 4-billion-parameter decision model built on Qwen3.5-4B-Base, with an adapter from developer Jared Palmer, under an Apache 2.0 licence.
The card lists three quantisations: a 4-bit file of 3.03 GB, an 8-bit file of 4.48 GB and a 16-bit file of 8.43 GB. It shows 245 downloads over the last month, and says no inference provider hosts it. We could not find release notes for the endpoint in the pages we opened, so how it behaves is unconfirmed.
The addition follows a week in which Perplexity and Cloudflare released their own small decision models. Whether the new endpoint works with those releases is not stated in the posts we read. No independent tests of Kev-4B or the endpoint have appeared yet, and its quality claims rest on its makers.