AI interpretability tools have a hidden flaw
A new audit finds that many features learned by AI interpretability tools are statistical artifacts with no real causal effect on model behavior.
A new audit finds that many features learned by AI interpretability tools are statistical artifacts with no real causal effect on model behavior.