English

PLUS ULTRAProduct LaunchesInco AISplash

Inco AI Unveils Splash Engine for Faster Local Qwen3.8 Execution on Apple Silicon

PLUS ULTRA by Amenoyomi

Inco AI has released Splash, an open-source inference engine designed specifically for running language models locally on Apple Silicon. The engine is optimized for Qwen3.6-35B-A3B and Qwen3.8-27B, providing dedicated GPU kernels and tailored memory plans for each supported model.

To improve generation speeds, each model is shipped with a dedicated DFlash 2 draft model to support speculative decoding, a technique that uses a smaller model to predict tokens to accelerate the inference process.

In internal tests conducted on a 48 GB M5 Pro, Splash achieved a decode speed of approximately twice that of the next-fastest engine measured for Qwen3.8-27B. Specifically, it delivered 74 tokens per second on short prompts and 54 tokens per second at 32K context. When handling four concurrent requests on short prompts, the combined throughput reached 170 tokens per second, which Inco AI reported as 3.9× the throughput of the next-fastest engine in their comparison.

Splash is available through LM Studio Bionic 1.1.5 or newer. The engine requires an M3 or newer Mac running macOS 26.4 or later with at least 36 GB of unified memory, though Inco AI recommends 48 GB or more. Users can install the engine by navigating to Settings > Runtime and selecting Splash (Metal) under the Experimental backends section.

PLUS ULTRAby Amenoyomi

To achieve these performance gains, Splash moves away from general-purpose inference paths by combining model-specific hardware architecture with predictive generation.

The engine employs GPU kernels and memory plans designed specifically for Qwen3.6-35B-A3B and Qwen3.8-27B. By tailoring the hardware interaction to the precise requirements of these specific models, Splash optimizes memory access and computation efficiency on Apple Silicon.

Generation speed is further increased through speculative decoding using dedicated DFlash 2 draft models. These smaller models predict upcoming tokens, allowing the main engine to verify and generate multiple tokens in a single step rather than processing them one by one.

This combination allows Splash to deliver roughly twice the decode speed of other engines for Qwen3.8-27B and reach up to 3.9 times the combined throughput during concurrent requests.

Sources

  1. Splash Engine - the fastest local Qwen3.8 on Apple Silicon (LM Studio Blog, 2026-09-18)