Safety gates for AI that hold under any generator
A new theory proves when AI safety filters stay reliable regardless of what model or policy is generating actions.
A new theory proves when AI safety filters stay reliable regardless of what model or policy is generating actions.