AI safety filters miss threats regex can't catch
LLM alignment alone blocks 0% of adversarial prompts by standard measures, but catches more when judged by another AI.
LLM alignment alone blocks 0% of adversarial prompts by standard measures, but catches more when judged by another AI.