One robot brain handles all four vision tasks
A single vision-language model now handles spatial, temporal, action, and state tasks for robots without losing specialized skills.
A single vision-language model now handles spatial, temporal, action, and state tasks for robots without losing specialized skills.