Tiny weight tweaks can gut AI safety
Changing less than 0.2% of a small AI model's weights can nearly disable its safety guardrails while keeping it otherwise useful.
Changing less than 0.2% of a small AI model's weights can nearly disable its safety guardrails while keeping it otherwise useful.