English

Product LaunchesSalvatore SanfilippoDwarfStar 4

DwarfStar 4: A Specialized Local Inference Engine for High-Memory Hardware

DwarfStar 4 is a narrow C inference engine specifically optimized for high-memory consumer hardware, including Mac (Metal), CUDA, and ROCm platforms. The engine is designed to facilitate the local execution of large Mixture-of-Experts (MoE) models by utilizing asymmetric quantization to target routed experts while preserving critical computational paths.

The project, created by Salvatore Sanfilippo, supports DeepSeek V4 and V4.1 Flash (including experimental vision models), GLM 5.x, and Qwen3.8 Flash Next. It provides a full stack that includes local APIs, a Command Line Interface (CLI), and a native agent. The engine is built to be self-contained and is specifically tuned for large models that typically require remote serving, making them practical on high-memory machines such as 128 GB laptops or larger workstations.

Key features include an integrated KV cache management system that allows saving long prefixes to SSD and resuming sessions via prompt hashes, avoiding the need for full re-prefilling. The software supports various deployment modes: ./ds4 for interactive chat, ./ds4-server for local API serving, and ./ds4-agent for persistent coding sessions. Additionally, it offers speculative decoding through the --mtp option for supported models and specialized "thinking" controls for reasoning-heavy models.

Sources

  1. From the creator of Redis; run LLM locally with ds4 (Hacker News Frontpage, 2026-10-02)
  2. GitHub