Skip to content
News briefMD-2026-0211

A quantisation format crosses libraries

Transformers can now load GGUF checkpoints directly, reusing llama.cpp's own kernels rather than reimplementing them.

2 minOpen-Source AI Models

Transformers has added support for running GGUF models, the quantised checkpoint format developed by the llama.cpp team. A checkpoint sized to fit a laptop's memory can now be loaded through the familiar interface rather than through a separate stack.

GGUF packages weights and metadata in a single file, including tokenizer information and an optional chat template, and supports a range of quantisation levels so a user can trade precision against memory footprint. Variants such as Q4_K_M mix precision across tensors rather than applying one setting uniformly.

The format is already the de facto standard for local inference. llama.cpp's engine powers Ollama, LM Studio and Jan; the team publishes quantised checkpoints on the Hub, and publishers including Unsloth, LM Studio Community and bartowski supply ready-made variants across quantisation levels. Hugging Face reports GGUF models have been downloaded millions of times.

The implementation choice is the interesting part. Rather than reimplementing the arithmetic, the work reuses llama.cpp's own ggml kernels through the kernels library, and reduces overhead in the generation loop. Initial focus is local inference on Apple Silicon, starting with the Qwen3.5 architecture.

Compatibility without performance would have been a checkbox. Borrowing the kernels that made the format fast in the first place is what makes the support usable, and it is a more honest engineering answer than a second implementation that runs at half the speed.

Retold from Hugging Face. This is a summary in our own words; follow the link for the original reporting.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined