One model, two roles
The same weights that predict images from a camera are also used to control a robot’s movements. That is an ambitious pairing: perception and action share a model, rather than being described as separate components.
Robot-training data is scarce, so Reka AI built an inverse dynamics model to extract control signals from ordinary videos found online. Rho-1 was trained for about three months on 320 H100 GPUs.
The harder question
Reka AI has worked on multimodal models before. In April 2024, it introduced Reka Core, a language model that handled different data types and competed with GPT-4, Claude 3 and Gemini Ultra on benchmarks. Rho-1 arrives amid broader researcher interest in world models.
I think the more consequential claim is not that one model can process several kinds of input, but that its visual predictions can also drive robot movement. The announcement gives a training recipe and describes the architecture, but says little about how well Rho-1 controls robots in practice. Without that evidence, the shared weights are a compelling design choice, not yet proof that one model can reliably bridge seeing and acting.
Source: the-decoder.com
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X