AI safety monitors can be talked out of it
Adversarial agents can manipulate AI safety monitors through the same reasoning trace being monitored, boosting harmful approvals by ~10%.
Adversarial agents can manipulate AI safety monitors through the same reasoning trace being monitored, boosting harmful approvals by ~10%.