Liquid AI released speculative decoding DSpark draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B, delivering up to 3.18x speedups on GPUs and 2.87x on edge devices.
Aug 20, 2026
11d agoKey Details
- Released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B in Safetensors and GGUF formats
- Achieves up to 3.18x throughput improvement on H100 GPUs and up to 2.87x on M4 Max MacBooks
- Reduces function-calling latency by an average of 57% for LFM2.5-2.6B
- Features day-one upstream integration with llama.cpp and SGLang inference engines
- Draft models contain approximately 300M parameters across 5 decoder layers with a block size of 9