Reka's Rho‑1 puts text, video and robot control in one 19B model
Reka says its 19-billion-parameter Rho-1 draws images, edits video and outputs robot joint trajectories from one network, trained on 320 H100 GPUs. It is a research preview, and Reka lists its own limits.
Reka AI has announced Rho-1, a 19-billion-parameter model that Reka says "understands and generates text, images, and video, reasons over them, and takes actions" inside a single neural network. The Decoder reported the launch on 5 October. Reka calls it a research preview, and the announcement gives a contact address for partnership inquiries rather than a download or a price.
Reka says text, pixels and physical actions are all tokens in one context window, so the model needs no tool calls or helper models. In its post, Rho-1 draws images, detects objects, animates scenes and edits video in one conversation. It can also stream video that responds to live instructions, which Reka describes as "infinite, steerable livestreams".
For robots, Reka says the model treats actions as native tokens and can simulate future movements and output joint trajectories without an external planning layer. It says it extracted control signals from ordinary internet video with an inverse dynamics model, because robot training data is scarce. The Decoder says no robot types or demonstrations were detailed.
Reka puts the base model at 0.79 times real time and a distilled version at about one second per video, and says it trained on 320 H100 GPUs over about three months. The post concedes structural drift in long sequences, inconsistent grounding across frames, brittle editing and a native resolution capped at 672 by 384. No outside benchmarks exist yet.