Developer splits a language model across seven microcontrollers
An open-source project runs a compressed 0.5-billion-parameter model across a cluster of ESP32-S3 chips wired together, with one board handling tokenization and six others splitting the transformer layers.
A developer split a small language model across seven ESP32-S3 microcontrollers wired together, distributing the neural network's layers the way a cluster of much larger machines would.
The project is posted to GitHub under the handle Low-Zi-Hong, and reached the Hacker News front page with 50 points and five comments. One microcontroller acts as a master, handling tokenization and embeddings. The other six process transformer layers in parallel, linked over an SPI daisy-chain, according to the repository.
The model is compressed to fit: a 0.4- to 0.5-billion-parameter language model using 1.58-bit BitNet ternary quantization, plus 4-bit compression for the embedding layer, with the resulting weights stored across the cluster's flash memory, the repository says.
The repository does not publish tokens-per-second figures or a hardware cost breakdown. It has drawn 88 stars and five forks on GitHub, and no independent benchmark of the setup has been published.