AI reasoning traces lie when it matters
On hard problems, AI models blindly follow corrupted reasoning steps; on easy ones, they ignore their own reasoning entirely.
On hard problems, AI models blindly follow corrupted reasoning steps; on easy ones, they ignore their own reasoning entirely.