AI agents misjudge safe and unsafe actions
A new benchmark reveals AI agents wrongly block safe actions 28x more often than they allow dangerous ones.
A new benchmark reveals AI agents wrongly block safe actions 28x more often than they allow dangerous ones.