RT-1: Robotics Transformer for Real-World Control at Scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Lisa Lee, Tsang-Wei Edward Lee, Sergey Levine, Yao Lu, Henryk Michalewski, Igor Mordatch, Karl Pertsch, Kanishka Rao, Krista Reymann, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Pierre Sermanet, Jaspiar Singh, Anikait Singh, Radu Soricut, Huong Tran, Vincent Vanhoucke, Quan Vuong, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Jialin Wu, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, Brianna Zitkovich · Google DeepMind, Google Research, Everyday Robots · 2023
Framework
TensorFlow
License
Apache-2.0
Stars
1,723
Summary
A large transformer model trained on 130k episodes of real robot manipulation to output discretized arm-and-gripper actions from RGB images and natural language instructions.
Abstract Summary
Key Points
- First large-scale transformer model trained end-to-end on real robot manipulation.
- 130K+ episodes across 13 months on Everyday Robots mobile manipulators.
- Pixel-to-action control with natural language instructions; no privileged state.
- Demonstrated scaling benefits and positive cross-robot transfer.
- Precursor to RT-2, Open X-Embodiment, and the broader VLA research wave.
Additional Notes
Training Tips
- Data diversity matters more than data volume alone; include many objects, lighting conditions, and backgrounds.
- Action token vocabulary size impacts coarse-vs-fine control precision.
- The model is sensitive to image preprocessing; matching the training resolution and normalization is critical.
Related Papers
- RT-2 (Brohan et al., 2023)
- Open X-Embodiment (Padalkar et al., 2023)
- Octo (Team et al., 2024)
Related Papers
ALAM: Algebraically Consistent Latent Transitions for Vision-Language-Action Models
Zuojin Tang, Haoyun Liu, Xinyuan Chang et al. · arXiv · May 2026
Vision-Language-Action (VLA) models map visual observations and language instructions to robot actions through learned l...
PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
Xinyu Guo, Bin Xie, Wei Chai et al. · arXiv · May 2026
Vision-Language-Action (VLA) models have shown promise for generalizable robot control but struggle when adapting to new...
Unified Noise Steering for Efficient Human-Guided VLA Adaptation
Junjie Lu, Xinyao Qin, Yuhua Jiang et al. · arXiv · May 2026
Adapting Vision-Language-Action (VLA) models to specific user preferences or task requirements typically requires expens...
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction
Kai Xiong, Hongjie Fang, Lixin Yang et al. · arXiv · May 2026
Effectively handling the interplay between spatial perception and action generation remains a critical bottleneck in robotic manipulation. Existing methods typically treat spatial perception and action execution as decoupled or strictly...