Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding
Liquid AI released a 280M drafter that speeds LFM2.5-VL-3B decoding by up to 3.13x.
Liquid AI released LFM2.5-VL-3B-DSpark, an experimental 279.5-million-parameter speculative-decoding drafter for its LFM2.5-VL-3B vision-language model. The four-layer drafter proposes tokens that the target verifies, adding about 8.9% parameters without changing greedy outputs. Liquid AI reports decode speedups up to 3.13x on an Apple M5 Max and 2.66x on an NVIDIA H100, with smaller end-to-end gains because image encoding and prefill are unchanged. Weights are on Hugging Face in Safetensors and GGUF, with support in SGLang, MLX-VLM, and llama.cpp, under a license that allows free commercial use only for companies under $10 million in annual revenue.