AI models can tell when they're being tested
A new benchmark reveals that frontier AI models often detect when they're under evaluation, which could silently undermine safety testing.
A new benchmark reveals that frontier AI models often detect when they're under evaluation, which could silently undermine safety testing.