
Up to 3.2x Faster Inference with LFM2.5-DSpark
Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B models. These utilize speculative decoding to increase inference speed without changing output quality.
Why it matters
This reduces the time users wait for AI responses on both GPUs and local devices. It specifically lowers latency for function-calling, improving how AI agents perform on personal hardware.
The details
- Inference is up to 3.18x faster on GPU and 2.87x faster on-device.
- Draft models are relatively small, each approximately 300M parameters.
- Includes day-one integration for llama.cpp and SGLang.
Show entities and relationshipsHide entities and relationships
In this article
Key connections
Liquid AI owns LFM2.5-DSpark
Liquid AI developed and released the LFM2.5-DSpark draft model checkpoints.
Liquid AI owns LFM2.5-8B-A1B
Liquid AI developed the LFM2.5-8B-A1B mixture-of-experts language model.
LFM2.5-DSpark uses Speculative Decoding
LFM2.5-DSpark uses speculative decoding to accelerate LLM inference.
LFM2.5-DSpark is related to LFM2.5-1.2B-Instruct
LFM2.5-DSpark includes draft model checkpoints for LFM2.5-1.2B-Instruct.
LFM2.5-DSpark is related to LFM2.5-2.6B
LFM2.5-DSpark includes draft model checkpoints for LFM2.5-2.6B.
LFM2.5-DSpark is related to LFM2.5-8B-A1B
LFM2.5-DSpark includes draft model checkpoints for LFM2.5-8B-A1B.
Show 6 more connectionsShow fewer connections
LFM2.5-DSpark uses llama.cpp
LFM2.5-DSpark draft models offer day-one support for on-device inference in llama.cpp.
LFM2.5-DSpark uses SGLang
LFM2.5-DSpark draft models offer day-one support for GPU inference in SGLang.
LFM2.5-DSpark uses H100
LFM2.5-DSpark draft models were benchmarked on a single H100 80GB GPU.
LFM2.5-DSpark uses MacBook Pro
LFM2.5-DSpark draft models were benchmarked on an M4 Max MacBook Pro.
LFM2.5-8B-A1B is built with Mixture of Experts
LFM2.5-8B-A1B is an MoE architecture.
LFM2.5-DSpark is built with DSpark
LFM2.5-DSpark is built using the DSpark speculative decoding draft architecture.
Related events
Liquid AI Releases DSpark Draft Model Checkpoints for LFM2.5 Family
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.