AI still fails real science tasks at scale
A new benchmark using real scientific requests found that no frontier AI model consistently meets expert-acceptable quality.
A new benchmark using real scientific requests found that no frontier AI model consistently meets expert-acceptable quality.