English

Model ReleasesCactus ComputeNeedle 3

Cactus Compute Releases Needle 3, a Foundation Model for Tiny Devices Featuring "Intelligence Laddering"

Cactus Compute has released Needle 3, a foundation model specifically designed for resource-constrained environments such as mobile phones, wearables, robots, smart homes, and microcontrollers. The model's size ranges from 8 MB to 29 MB in 2-bit quantized format.

Needle 3 employs a "Laddered Simple Attention Network" architecture. This approach creates an "intelligence ladder" where every depth from 2 to 20 layers functions as a standalone, deployable sub-network. This allows developers to choose a sub-network size tailored to their specific hardware constraints. According to the developers, a 4-layer sub-network fine-tuned on downstream tasks can match the performance of DeepSeek V4 Flash, despite starting at only 29 million parameters.

The model is optimized for automation tasks, specifically tool calling and structured data extraction. It uses a 2-bit quantization and a proprietary Simple Attention Network to trade general chat capacity for high accuracy in function calling and text extraction on small devices. When fine-tuned for specific tasks, such as the "DroidCall" dataset, sub-networks showed significant performance improvements of 18 to 36 points.

For implementation, Cactus Compute provides a Python package and an inference engine under 1 MB. The engine can be built for various platforms, including Linux-arm64 and macOS-arm64, to support a wide range of deployment targets.

Sources

  1. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash (Hacker News Frontpage, 2026-09-18)
  2. GitHub