AI safety tests may measure capability, not safety
A study finds popular agent-safety benchmarks often measure raw capability rather than genuine safe behavior.
A study finds popular agent-safety benchmarks often measure raw capability rather than genuine safe behavior.