Embodied-R1 points instead of acting and hits 87.5% on real robot tasks
Robots increasingly see the world through a camera and read our written instructions. But that "knowledge" often fails to turn into the right action: the model knows what a cup is, yet not where to put it or how to get around the objects next to it. This distance between vision and action is the seeing-to-doing gap. The Embodied-R1 team proposes…