Up to 3.2x Faster Inference with LFM2.5-DSpark

LiquidAI releases DSpark draft model checkpoints for LFM2.5 variants to accelerate inference via speculative decoding. The update claims up to 3.18x throughput gains on GPU and integrates with llama.cpp and SGLang. Developers can use these open-sourced weights to reduce latency without altering output quality.

Cover image for Up to 3.2x Faster Inference with LFM2.5-DSpark