Granite 4.2 adds a switch for thinking
IBM ships 3B, 8B and 30B under Apache 2.0, with chain-of-thought that can be turned off and a low-effort mode for simple queries.
3 minOpen-Source AI Models
IBM published Granite 4.2 on 25 August: three dense decoder-only models at 3B, 8B and 30B parameters, Apache 2.0, on Hugging Face and GitHub with FP8, FP4 and GGUF variants. All three use grouped query attention with 40 heads and 8 KV heads, rotary embeddings, SwiGLU and RMSNorm, and carry a 131,072-token maximum sequence length, extended to 512K during the fifth pre-training phase.
What is new against 4.1
Reasoning. The models produce explicit chain-of-thought, and it can be switched off. There is a low-effort thinking setting for simple queries, and the 8B and 30B were given agentic reinforcement learning in real sandboxed environments covering software engineering, terminal use and search.
The toggle is the part worth noting. A model that always reasons is a model that always bills for reasoning, and the failure mode this desk documented on Qwen 3.8 last week, 22,276 reasoning tokens on a request that needed almost none, is a cost problem before it is a quality one. Making it a runtime setting rather than a prompt convention puts the control where an operator can enforce it.
The training pipeline
- Pre-training on about 15 trillion tokens across five phases.
- Supervised fine-tuning on 7.2 million samples, roughly 100B tokens, split 31.6% agentic and 68.4% non-agentic.
- Post-training with GRPO: foundational RL on maths, coding and reasoning, then agentic RL for the two larger models, then RLHF alignment.
Reported scores
- AIME25: 78.33% at 3B rising to 89.17% at 30B.
- SWE-Bench Verified: 47.67% at 8B, 57% at 30B.
- MMLU-Pro: 67.84% to 77.60% across the three sizes.
- Twelve languages, English, German, Spanish and Chinese among them.
Retold from Hugging Face. This is a summary in our own words; follow the link for the original reporting.