Hidden AI safety gaps found inside models
A new analysis finds that AI safety training can look effective from the outside while the model's internal state still harbors dangerous content.
A new analysis finds that AI safety training can look effective from the outside while the model's internal state still harbors dangerous content.