Skip to content
BenchmarkMD-2026-0209

Three hundred million parameters doing the guessing

Liquid AI's DSpark draft models are around 300M parameters and five decoder layers, and report up to 3.18× throughput on an H100 and 2.87× on a MacBook.

3 minOpen-Source AI Models

Liquid AI published DSpark on 20 August: speculative decoding for the LFM2.5 family. A small draft model proposes candidate tokens, the target model verifies them all in one forward pass, and anything the target rejects is discarded. The output is identical to what the target would have produced alone; only the number of forward passes changes.

The drafters are small. LFM2.5-1.2B-Instruct-DSpark is 295.7M parameters; the drafters for LFM2.5-2.6B and LFM2.5-8B-A1B are 327.7M each. All three are attention-only with five decoder layers.

The measured numbers

  • H100 — LFM2.5-2.6B: 323 → 864 tok/s, 2.67× average. LFM2.5-1.2B-Instruct: 656 → 1,384 tok/s, 2.10×. LFM2.5-8B-A1B: 418 → 1,074 tok/s, 2.54×. Peak reported: 3.18×.
  • M4 Max MacBook — LFM2.5-1.2B-Instruct: 138 → 350 tok/s, 2.54×. LFM2.5-2.6B: 61 → 139 tok/s, 2.27×. LFM2.5-8B-A1B: 90 → 106 tok/s, 1.18×. Peak reported: 2.87×.
  • Function calling on LFM2.5-2.6B: 57% lower latency.

The row that tells you most is the 8B mixture on the MacBook: 1.18×. Speculative decoding pays when verification is cheap relative to generation and the draft guesses well. On a memory-bandwidth-bound device running a model with one billion active parameters, neither condition holds strongly, and the technique nearly disappears. The same architecture on an H100 returns 2.54×.

Why function calling gains most

Structured output is the easiest thing in the world to guess. Field names, brackets, quotes and separators are close to determined by what came before, so the draft model's acceptance rate is high exactly where latency is most visible to a user waiting on a tool call. A 57% reduction on that path is a larger practical result than the throughput headline.

Availability

Safetensors and GGUF on Hugging Face, with day-one support in llama.cpp and SGLang, both open-sourced. Anyone already serving LFM2.5 can measure this against their own traffic rather than against the published mix — and given the spread between the H100 and MacBook rows, that is the only measurement worth acting on.

Retold from Hugging Face. This is a summary in our own words; follow the link for the original reporting.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined