English

Product LaunchesCactus

Cactus Releases Open-Source AI Inference Engine Optimized for Smartphones

This article is a translation. Read the Japanese original

Cactus is an inference engine designed specifically for the constraints of mobile devices, with a strong focus on reducing latency, protecting privacy, and enabling offline operation.

The team cited a lack of support for low-to-mid-range devices and high battery consumption as primary issues with existing frameworks. To address this, Cactus implements its own kernels to optimize energy efficiency and accelerator support.

Regarding performance, the team provided CPU benchmarks for the Qwen3-600m-INT8 model. The Pixel 9, Galaxy S25, and iPhone 16 recorded speeds of 50 to 70 tokens per second. In NPU environments, Qwen3-4B-INT4 operates at 21 tokens per second.

For developers, an SDK is provided that allows app developers to build agent workflows with just 2 to 5 lines of code. Apps such as AnythingLLM and KinAI are already using it in production, processing over 500,000 inference tasks per week.

The license is free for hobbies and personal projects, while a paid license is required for commercial use. The code is available on GitHub.


Source: Launch HN: Cactus (YC S25) – AI inference on smartphones(HN 123pt・63コメント) (HN Search (backfill), 2025-09-19)