Octo: An Open-Source Generalist Robot Policy
FeaturedDibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, Sergey Levine · UC Berkeley, Stanford, TRI · 2024
Framework
JAX / Flax
License
MIT
Stars
1,652
Summary
Octo is a large open-source transformer-based generalist robot policy trained on 800k trajectories, supporting language-conditioned and goal-image-conditioned control across 9 robotic platforms with efficient fine-tuning on consumer GPUs.
Abstract Summary
Key Points
- Large open-source generalist policy trained on 800k trajectories from Open X-Embodiment.
- Supports language and goal-image conditioning with a flexible transformer architecture.
- Observation-action tokenizer enables cross-robot fine-tuning on consumer GPUs within hours.
- Evaluated on 9 distinct robot platforms with strong few-shot transfer to new tasks.
- Full training and fine-tuning stack released in JAX/Flax with Hugging Face integration.
Additional Notes
Fine-tuning Tips
- Use the provided observation tokenizer config to map your robot’s cameras to the pretrained input space.
- Fine-tuning works best with 50–500 in-domain trajectories; less data often leads to overfitting.
- The JAX ecosystem benefits from TPUs for large-scale pre-training, but GPU fine-tuning is well supported.
Related Papers
- Open X-Embodiment (Padalkar et al., 2023)
- RT-2 (Brohan et al., 2023)
- OpenVLA (Kim et al., 2024)
Related Papers
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction
Kai Xiong, Hongjie Fang, Lixin Yang et al. · arXiv · May 2026
Effectively handling the interplay between spatial perception and action generation remains a critical bottleneck in robotic manipulation. Existing methods typically treat spatial perception and action execution as decoupled or strictly...
DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset
Alexander Khazatsky, Karl Pertsch, Suraj Nair et al. · RSS 2024 · Jul 2024
A 350-hour dataset of diverse real-world robot manipulation across 22 robots in 71 scenes, designed to train scalable and generalist imitation learning policies.
LeRobot: A Library for Real-World Robot Learning
Remi Cadene, Simon Alibert, Alexander Soare et al. · NeurIPS 2024 Workshop · Jun 2024
Hugging Face's LeRobot is an open-source PyTorch framework providing pretrained models, datasets, and training scripts for imitation and reinforcement learning on real robots, lowering the entry barrier to robot learning.
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Open X-Embodiment Collaboration · arXiv · Oct 2023
The largest collaborative robot learning dataset ever assembled, spanning 22 robots from 21 institutions, enabling training of generalist RT-X policies that exhibit positive cross-embodiment transfer.