AI lab robots ace benchmarks but fail new tasks
LLM agents controlling microscopes pass known tests but can't reliably generalize to unfamiliar tasks.
LLM agents controlling microscopes pass known tests but can't reliably generalize to unfamiliar tasks.