Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

https://arxiv.org/abs/2606.00206

Comments

DiabloD3Sep 28, 2026, 1:37 AM
The title of the paper is correct. The paper does not seem to actually get to the point in a generic way, but hyperfocuses on, effectively, one type of error compensation.

Highly quantized models, especially with highly quantized KV caches, will, effectively, attend to the wrong tokens and be unable to easily discern highly similar tokens. The bastardized way of explaining this is gradient descent techniques get stuck in localized minimum and global maximums, so what happens when you turn the slopes into hard stair steps?

We need to move to smaller models and smaller caches and better samplers, not new quant methods (although I'm willing to also take those too).