slotstream has been released as a tool for running the 125B parameter Qwen3.8-Flash-Next with 4-bit quantization on low-memory Macs. With a Mac-native implementation using MLX and Swift, it reportedly allows models that typically require over 100GB of memory to operate on as little as 16GB of memory through expert offloading and SSD streaming.

It is reported to achieve an inference speed of approximately 12 tok/s on a 48GB Mac. The tool includes an "auto-mode," which provides a function to automatically balance memory usage and speed. Installation and updates are also said to be easily performed.

As for the next steps, the developer plans to implement and port the MTP module for speculative decoding.


Source: Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s (Hacker News Frontpage, 2026-09-02)