Liquid AI Debuts DSpark Drafter to Accelerate LFM2.5-VL-3B Vision-Language Model Decoding

Liquid AI released LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The roughly 279.5M-parameter drafter reportedly speeds decoding by up to 3.13x on Apple silicon and 2.66x on NVIDIA H100 while preserving greedy-decoding outputs. Weights are available on Hugging Face with support in SGLang, MLX-VLM, and llama.cpp under the LFM Open License v1.0.
Liquid AI’s experimental drafter adds about 279.5M parameters to LFM2.5-VL-3B. Its four-layer attention stack includes hidden-state projection, a Markov head, and confidence components. The embedding and language-model head stay tied to the target, lifting deployed parameters by roughly 8.9%. Training used supervised fine-tuning data for common vision-language tasks over ten epochs on AMD hardware.
The drafter reads hidden states from several target layers and predicts multiple upcoming tokens, which the larger model checks together. Since text and image patches are tensors by those layers, one inference method serves both. Liquid AI suggests block sizes of eight or nine by hardware and reports gains across VQA, captioning, chart, reasoning, and conversation benchmarks.
Faster vision-language decoding could make multimodal assistants more responsive on laptops and phones, potentially lowering inference costs for developers and researchers. Smaller companies may benefit most, since the license permits free commercial use below $10M in annual revenue. Yet the drafter is experimental, and gains vary by hardware, workload, and concurrency, so real-world adoption may remain limited until broader validation and tooling mature.