Top-ranked AI fails half of real tasks
DeepSeek's top-benchmarked model completed only 54% of real-world agent tasks across Gmail, GitHub, and Slack.
DeepSeek's top-benchmarked model completed only 54% of real-world agent tasks across Gmail, GitHub, and Slack.