ESP32-S3 cluster running a 1.58-bit (BitNet) language model
This project splits a small language model across seven ESP32-S3 boards, passing intermediate activations along a daisy-chained SPI link. The master handles tokenization, embeddings, and sampling; six nodes each run part of the transformer layers. It is an unusually concrete embedded/distributed-AI experiment, with firmware, model preparation tools, and wiring instructions in the repository.
Treat the repository’s model-size and architecture description as a project report, not an independently benchmarked result: it does not yet provide inference speed or quality measurements. In the HN discussion, readers were intrigued by the build but questioned whether distributing inference over many microcontrollers can overcome communication and memory-bandwidth costs.