Liquid AI released an experimental version of its vision-language model, LFM2.5-VL-3B-DSpark, on September 24, 2026. The model applies "DSpark," a speculative decoding technology, to the existing LFM2.5-VL-3B to achieve significantly faster inference speeds without compromising output quality.
Model ReleasesLiquid AILFM2.5-VL-3B-DSpark
Liquid AI Releases LFM2.5-VL-3B-DSpark to Accelerate Vision-Language Model Inference
Speculative decoding works by using a lightweight "draft model" to predict output tokens in advance, which the larger main model then verifies. In this implementation, the draft model consists of approximately 280 million parameters, increasing the total parameter count by 8.9% and requiring more memory, but the performance gains are substantial.
On edge devices such as the M5 Max MacBook Pro, the model achieved up to a 2.62x increase in end-to-end throughput and up to a 3.13x increase in decoding speed. On NVIDIA H100 GPUs, end-to-end throughput improved by up to 2.27x, with decoding speed increasing by up to 2.66x.
The vision-specific drafter was developed by applying the same algorithm used for Liquid AI's text-based foundation models. Because the drafter operates on multi-dimensional tensors rather than raw input modalities, the algorithm is applicable to both text and image-based workloads.
LFM2.5-VL-3B-DSpark is available as an open model on Hugging Face in Safetensors and GGUF formats. It is already supported by inference ecosystems including llama.cpp, MLX-VLM, and SGLang.
Sources
- 軽量な視覚言語モデル「LFM2.5-VL-3B」にDSparkのドラフトモデルを適用して爆速化した「LFM2.5-VL-3B-DSpark」が登場 (GIGAZINE, 2026-09-25)
- Liquid AI Blog