One toolkit to run AI agents on 80+ benchmarks
A new open-source toolkit unifies testing for AI agents across 80+ benchmarks, with a curated hard task set no model yet passes 30% of.
A new open-source toolkit unifies testing for AI agents across 80+ benchmarks, with a curated hard task set no model yet passes 30% of.