AI safety tests can't prove what we think
A new framework shows exactly which safety claims AI red-team tests can and cannot support — and most miss rare, catastrophic risks by orders of magnitude.
A new framework shows exactly which safety claims AI red-team tests can and cannot support — and most miss rare, catastrophic risks by orders of magnitude.