AI safety that tracks intent across a conversation
A new framework tracks the full arc of a conversation to judge whether a user's intent is harmful, making AI harder to manipulate gradually.
A new framework tracks the full arc of a conversation to judge whether a user's intent is harmful, making AI harder to manipulate gradually.