Qwen 3.8 27B thinks too hard by default
Apache 2, 27 billion parameters, runs on a laptop — and spends 22,276 reasoning tokens on a request that needed almost none.
2 minOpen-Source AI Models
Qwen 3.8 27B was released on 15 August under Apache 2, with vision capability and a quantised build of about 17GB. Simon Willison ran it on a 128GB M5 Max MacBook Pro and on an NVIDIA DGX Spark, both through LM Studio, at 15 to 30 tokens per second — roughly 72% faster with multi-token prediction enabled. Context was tested up to 262,144 tokens.
The default that costs you
The model ships with reasoning effort set to xhigh, and the consequence is measurable. One SVG generation consumed 22,276 reasoning tokens to produce 3,223 tokens of output, taking 21 minutes. Asked simply to draw a circle, it returned an elaborately animated one after minutes of deliberation — a good answer to a question nobody asked.
The recommendation that follows is blunt: run it with low reasoning, or none, for ordinary work. The default adds latency without a matching gain on straightforward requests.
Where it did well is worth recording too. Bounding-box detection matched closely on a 0-1000 scale, it built working tools and Python utilities, and it ran inside a coding agent framework without trouble. SVG output was better than smaller local models produce.
For a locally-run open-weight model the interesting number is not the benchmark but the token budget. A model that thinks for 22,000 tokens before answering is cheap in licence terms and expensive in everything else.
Retold from Simon Willison. This is a summary in our own words; follow the link for the original reporting.