AI explanations fail regulatory stress tests
Mechanistic interpretability findings contradict each other 73% of the time under varied analysis settings, making them unreliable for legal compliance.
Mechanistic interpretability findings contradict each other 73% of the time under varied analysis settings, making them unreliable for legal compliance.