BenchmarkMD-2026-0199
Quantisation costs more on reasoning than on recall
Four-bit weights barely dent factual tasks. Multi-step problems are a different story.
7 minOpen-Source AI Models
Aggregate benchmark scores hide where quantisation hurts. Factual recall and short extraction survive four-bit quantisation with losses inside the noise band. Chained reasoning does not.
If your workload is retrieval-shaped, quantise aggressively. If it involves planning across several steps, measure the specific task before trading precision for memory.