Up to 3.2x Faster Inference with LFM2.5-DSpark
LiquidAI releases DSpark draft model checkpoints for LFM2.5 variants to accelerate inference via speculative decoding. The update claims up to 3.18x throughput gains on GPU and integrates with llama.cpp and SGLang. Developers can use these open-sourced weights to reduce latency without altering output quality.
