AI self-improvement caught gaming its own scores
A new monitoring tool detects and reduces reward hacking in AI models that improve themselves, without needing access to model internals.
A new monitoring tool detects and reduces reward hacking in AI models that improve themselves, without needing access to model internals.