Magnitude, an open‑source engine that tunes itself to your hardware
Magnitude says it compiles kernels on your own machine and runs up to 2x faster than llama.cpp. Its Apache 2.0 repo has 5,600 stars.
Magnitude, a Y Combinator S25 startup, launched an open-source inference engine for running open-weight models locally, and its Launch HN post reached 80 points and 40 comments on the Hacker News front page on Wednesday. Its GitHub repo showed about 5,600 stars and 396 forks when checked.
According to the README, the engine compiles and tunes kernels on the user's own device before running a model, rather than shipping pre-compiled ones. The developers say this lets it adapt to exact hardware, and that it is licensed under Apache 2.0. It targets Apple Silicon, NVIDIA GPUs, AMD GPUs and CPU-only machines.
The project's own benchmarks claim up to 2x the speed of llama.cpp. It lists 92% faster decode on Metal and 19% faster on CUDA, with prefill gains of 9% on Metal and 23% on CUDA. These numbers are the developers' and no outside group has reproduced them.
Magnitude also says it uses 27% less memory per agent by sharing prefix caches across concurrent sessions. The pitch is aimed at coding agents that run several sessions at once against one local model, which could matter on laptops.
Setup, according to the README, means downloading a desktop app from magnitude.dev, choosing a model and connecting an agent such as Pi, OpenCode or Hermes in one click. A command line tool comes bundled. The comparison settings and the models tested are the developers' choice.