ByteShape has released the full ShapeLearn models for Qwen 3.8 27B, providing a new set of GGUF-quantized versions that utilize hybrid quantization techniques to maintain high quality at low bitlengths. These models replace the previously released ShapeLearn-Lite versions.
Model ReleasesByteShapeQwen3.8-27BShapeLearn
ByteShape Releases ShapeLearn Quantized Models for Qwen 3.8 27B
The release includes various configurations designed to balance throughput and model accuracy. In benchmarking across six different GPUs, the ShapeLearn models consistently appeared on the performance frontier, meaning no other tested models were both faster and more accurate. For example, on an NVIDIA RTX 5090, the GPU-5 configuration achieved 93.7 tokens per second while maintaining 99.63% of the BF16 baseline's aggregate benchmark score.
The models also support speculative decoding to increase throughput. Users can utilize either the embedded MTP (Multi-Token Prediction) head, which adds approximately 250 MB to the file, or the external DFlash2 draft model from Inco AI, which requires roughly 1.1 GB of additional VRAM. Testing showed that DFlash2 provided higher throughput gains than MTP in most cases, reaching 1.34–2.10x the baseline next-token prediction throughput.
Sources
- Shapelearn Qwen 3.8 27B (13.1 GB VRAM) (Hacker News Frontpage, 2026-09-18)
- Hugging Face