Safety built into AI, not bolted on
Tiny classifiers reading an AI's internal states can catch harmful outputs before they happen, matching larger safety tools at lower cost.
Tiny classifiers reading an AI's internal states can catch harmful outputs before they happen, matching larger safety tools at lower cost.