Browser tool lets users run seven tiny LLMs
MicroLLM Lab runs 25-million to 360-million-parameter models entirely on-device via WebGPU, compressing them to 4 bits per parameter so they fit in under 85 megabytes of memory.
A browser-based tool called MicroLLM Lab lets users load and compare small language models, according to the project's own page. The models range from 25 million to 360 million parameters. Nothing needs installing, and no account is required. The project reached the Hacker News front page Monday with 88 points and 32 comments.
The site runs models locally using WebGPU, with WebAssembly and plain JavaScript as fallbacks, the page says. It compresses each model from 16-bit floating point down to 4 bits per parameter, cutting memory use by about 75%. That leaves the largest models needing roughly 50 to 84 megabytes of browser memory. The page reports sub-10-millisecond time to a model's first output token.
Models are cached in the browser's own IndexedDB storage once downloaded, the page says. No prompt or model data is sent to a server after the first load, it adds. The project credits an existing open project, petitgpt, as its inspiration. Its own page does not name who built it beyond that credit.
The models on offer are far smaller than mainstream chatbots. They are better suited to showing how local inference works than to serious daily use.