AI reasoning exposes hidden model secrets
Amplifying a model's 'thinking out loud' tendency can surface hidden behaviors up to 10x more often than normal auditing.
Amplifying a model's 'thinking out loud' tendency can surface hidden behaviors up to 10x more often than normal auditing.