Liquid AI has released an experimental DSpark draft model designed to accelerate inference for its vision-language model (VLM), LFM2.5-VL-3B. The draft model utilizes speculative decoding, a technique that employs a smaller, faster model to propose candidate tokens, which the larger target model then verifies. This approach provides a significant speedup with only a minimal increase in memory footprint.
Model ReleasesLiquid AIDSparkLFM2.5-VL-3B
Liquid AI Releases DSpark Draft Model to Accelerate LFM2.5-VL-3B Inference
The DSpark vision drafter utilizes the same architecture as the text-based LFM2.5-DSpark models. It captures hidden states from a fixed set of layers in the target model and operates on hidden-state vectors of identical dimensionality regardless of the input modality. The resulting drafter contains approximately 280 million parameters, increasing the deployed model's parameter count by 8.9%.
Performance measurements indicate substantial improvements across various hardware configurations. In on-device inference using Apple silicon, decoding speeds increased by 2.30x to 3.13x on an M5 Max and by 1.57x to 2.14x on an M3 Ultra. For GPU inference on an NVIDIA H100, the drafter delivered 20.4x to 2.66x faster decoding.
The LFM2.5-VL-3B DSpark draft model includes day-one support for llama.cpp, MLX-VLM, and SGLang. The model is available on Hugging Face in Safetensors and GGUF formats.
Sources
- Accelerating vision-language models with LFM2.5-VL-DSpark (Hugging Face Blog, 2026-09-24)