Multi-turn attacks broke AI safety 88% of the time
Attackers who adapted across a conversation bypassed AI safety measures 88% of the time — single-turn testing missed it entirely.
Attackers who adapted across a conversation bypassed AI safety measures 88% of the time — single-turn testing missed it entirely.