Needle 3: 8–29 MB automation models for tiny devices
Cactus’s Needle 3 targets tool calls and structured JSON rather than open-ended chat: deployable 2-bit subnetworks range from 25M to 121M parameters in 8–29 MB binaries, with reported Raspberry Pi 5 decode rates up to 4,000 tokens/s. The architecture uses a Monarch/Hadamard-factorized MLP, multilingual support, confidence scores, and runs across Linux, Android, iOS, RISC-V, MIPS32, microcontrollers, and browsers.
The important boundary is task scope: HN testers found the demo brittle on ambiguous smart-home commands, while the authors emphasize that the model is for narrow, grounded automation on hardware that cannot host a multi-billion-parameter model. That makes it interesting embedded inference, not a tiny general-purpose LLM.