The Reasoning Model Factuality Paradox: Why the Smartest LLMs Are the Least Reliable on Basic Facts
HalluHard 2026 and a year of production telemetry now agree on something uncomfortable — the same chain-of-thought that makes frontier models better at hard problems makes them measurably worse at staying faithful to the documents in front of them. This is the central deployment problem of 2026.