Needle is an experimental model designed to operate on consumer devices such as smartphones and smartwatches. Cactus announced that the model achieves speeds of 6,000 tok/s for prefill and 1,200 tok/s for decoding on consumer hardware.
Cactus Releases Needle, a 26M Function Calling Model Distilled from Gemini
This article is a translation. Read the Japanese original
The architecture employs "Simple Attention Networks," which completely eliminates MLPs and consists solely of attention and gating. This design is based on the view that tool calling is not a matter of inference, but rather a process of "search and assembly," such as matching queries with tool names and extracting arguments [Hacker News].
For training, the developers performed 27 hours of pre-training on 200B tokens using 16 TPU v6e units. This was followed by 45 minutes of post-training using 2B tokens of synthetic data generated by Gemini, covering 15 tool categories including timers, messaging, and navigation.
The developers stated that they gained the insight that FFNs (Feed-Forward Networks) are unnecessary for tasks involving access to external structured knowledge. They reported that in single-shot function calling, Needle outperformed models such as FunctionGemma-270M and Qwen-0.6B [Hacker News].
This project is part of Cactus, an inference engine for mobile and wearable devices. The model weights are available on Hugging Face, and the code is released under the MIT license on GitHub.
Source: Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model(HN 776pt・211コメント) (HN Search (backfill), 2026-05-13)