Accelerating vision-language models with LFM2.5-VL-DSpark

LiquidAI released a 280M parameter draft model to accelerate LFM2.5-VL-3B inference. The speculative decoding component claims speedups of up to 3.13x on local hardware and 2.66x on H100 GPUs. Day-one integrations are available for llama.cpp, MLX-VLM, and SGLang. The additional memory overhead is reported at 8.9% of the target model size.

Cover image for Accelerating vision-language models with LFM2.5-VL-DSpark

LiquidAI released an experimental draft model for its LFM2.5-VL-3B vision-language architecture. This twenty-eight million parameter component uses speculative decoding to accelerate inference on local hardware and server-grade GPUs. The team reports that this addition increases the total memory footprint by only eleven percent of the target model size. Support for this workflow is currently available through three distinct software backends for developers. The primary benefit is a significant reduction in inference latency without altering the quality of the generated output. On specific Apple silicon chipsets, decoding speeds improved by a factor of three point one three. When running on H series hardware, the speedup reaches two point six six times the baseline performance. These gains allow larger models to operate efficiently on edge devices rather than exclusively in data centers. The release notes acknowledge that this acceleration technique does not speed up the vision encoding phase. Consequently, overall system latency may still be constrained by tasks where image processing dominates the compute time. The authors apply Amdahl's law to explain why individual component speedups do not always translate to proportional total gains. Specific performance variances across different model tasks remain a limiting factor for consistent throughput.