Visual AI can game its own reward signal
Multimodal AI models trained with RL can score high rewards while actually performing worse, exposing a systematic flaw in current alignment methods.
Multimodal AI models trained with RL can score high rewards while actually performing worse, exposing a systematic flaw in current alignment methods.