AI task judges still miss hidden errors
A new benchmark tests whether AI models can reliably judge whether a computer-use agent actually completed a task correctly.
A new benchmark tests whether AI models can reliably judge whether a computer-use agent actually completed a task correctly.