Recent
- News briefMD-2026-0211
A quantisation format crosses libraries
Transformers can now load GGUF checkpoints directly, reusing llama.cpp's own kernels rather than reimplementing them.
2 minSource: Hugging Face
- News briefMD-2026-0210
Granite 4.2 adds a switch for thinking
IBM ships 3B, 8B and 30B under Apache 2.0, with chain-of-thought that can be turned off and a low-effort mode for simple queries.
3 minSource: Hugging Face
- BenchmarkMD-2026-0209
Three hundred million parameters doing the guessing
Liquid AI's DSpark draft models are around 300M parameters and five decoder layers, and report up to 3.18× throughput on an H100 and 2.87× on a MacBook.
3 minSource: Hugging Face
- BenchmarkMD-2026-0208
Recovering most of what quantisation costs
Liquid AI's Q4_0 checkpoints hold about 97% of full-precision accuracy while decoding 3–33% faster than the quantisations they replace.
2 minSource: Hugging Face
- Field reportMD-2026-0207
Qwen 3.8 27B thinks too hard by default
Apache 2, 27 billion parameters, runs on a laptop — and spends 22,276 reasoning tokens on a request that needed almost none.
2 minSource: Simon Willison
- News briefMD-2026-0206
A 3B vision model that fits in 3GB
Liquid AI's LFM2.5-VL-3B targets screen understanding on device, and publishes speed figures for CPUs rather than only for datacentre parts.
2 minSource: Hugging Face
- AnalysisMD-2026-0205
The licence is the specification
Open weights are not open source, and the difference decides what you are allowed to ship.
10 min
- BenchmarkMD-2026-0199
Quantisation costs more on reasoning than on recall
Four-bit weights barely dent factual tasks. Multi-step problems are a different story.
7 min
- Field reportMD-2026-0193
Checkpoint provenance is mostly unverified
Teams pull weights from public hubs and rarely check what they received.
6 min
- AnalysisMD-2026-0188
Fine-tuning is back for narrow tasks
Prompting won the general case. For classification and extraction the economics reversed again.
8 min
- Research noteMD-2026-0182
Model cards have stopped being useful
The format survived. The candour did not.
5 min