PrismML, based in the United States, has announced "Ternary Bonsai 2 27B," a model based on Alibaba's "Qwen3.8 27B" that achieves high local throughput and excellent energy efficiency.
PLUS ULTRAModel ReleasesPrismMLTernary Bonsai 2 27B
PrismML Announces "Ternary Bonsai 2 27B," Shrinking Qwen3.8 27B to 5.9GB
PLUS ULTRA by Amenoyomi
This article is a translation. Read the Japanese original
This model reduces its size to 5.9GB—just one-ninth the size of the full-precision model—while maintaining 98.2% of its performance in aggregate benchmarks. The size makes it viable for local execution even on graphics cards with only 8GB of VRAM.
Ternary Bonsai 2 27B reportedly achieves improvements in inference, coding, vision, and long-term agent performance. It is also characterized by its ability to run on Apple devices, including Mac, iPhone, and iPad.
The model is released under the Apache 2.0 license and is also available on Hugging Face.
PLUS ULTRAby Amenoyomi
The compression efficiency, which condensed the model to 5.9GB—one-ninth of the full-precision size—while maintaining 98.2% of its performance, represents an extremely high standard in AI model distribution.
By shrinking the model to this size, it becomes a realistic option to run a 27-billion-parameter class large-scale model in a local environment, even on mainstream graphics cards equipped with only 8GB of VRAM. Furthermore, compatibility with Apple devices such as iPhone, iPad, and Mac significantly lowers the barrier to entry caused by hardware constraints.
This version evolves beyond simple weight reduction by refreshing the base model from Qwen3.6-27B to Qwen3.8 27B compared to the previous generation, Bonsai 27B. With improvements in inference, coding, vision, and long-term agent performance, it has evolved into a compact model with practical, high-level performance.
Sources
- Qwen3.8 27Bを5.9GBまで小型化しつつ性能は98.2%維持したAIモデル「Ternary Bonsai 2 27B」が登場 (GIGAZINE、2026-09-18)