A 280M drafter for a 3B vision model
Liquid AI's draft model for LFM2.5-VL-3B decodes up to 3.13x faster on an M5 Max; end-to-end gains are smaller because prefill is untouched.
2 minOpen-Source AI ModelsFresh · 24 Sept
Liquid AI has released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model. Like the company's earlier text drafters, it adds a speculative decoding path: a small model proposes a block of tokens and the 3B target verifies them. Because the target checks every proposed token, greedy output is identical to running the target alone. The gain is speed, not different answers.
The drafter is an attention-only stack of four layers with a block size of nine, chosen after ablations across three, four and five layers. It has about 280M parameters, 193.0M of them in the decoder stack, 65.5M in a Markov head and 21.0M in a hidden-state projection, and it adds 8.9% to the deployed model. Image patches and text tokens are projected into the same representation before the layers the drafter reads, so the algorithm is unchanged from the text models. On six vision tasks following the MMSpec benchmark, with a block size of eight, the speedups were:
- MLX on an M5 Max: decoding 2.30x to 3.13x faster, end-to-end latency 1.56x to 2.62x better.
- llama.cpp on an M3 Ultra: decoding 1.57x to 2.14x faster, end-to-end 1.30x to 1.77x.
- H100: decoding up to 2.66x faster, end-to-end 1.64x to 2.27x.
The gap between decode and end-to-end is the most useful part of the post. Speculative decoding accelerates only the decode stage; the vision encoder and the prefill over hundreds of visual tokens run as before. On edge hardware, where prefill takes a larger share of wall time, even a large decode speedup yields a modest overall gain, which the authors describe as Amdahl's law at work.
Support is available from day one in llama.cpp, MLX-VLM and SGLang, each through a specific build or pull request, and the weights are published on Hugging Face in Safetensors and GGUF formats.
Retold from Hugging Face. This is a summary in our own words; follow the link for the original reporting.