AI learns safe behavior from human reasons
A new method trains AI agents to behave safely by learning from human preferences and written explanations, without needing a reward function.
A new method trains AI agents to behave safely by learning from human preferences and written explanations, without needing a reward function.