ACT: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
FeaturedTony Z. Zhao, Vikash Kumar, Sergey Levine, Chelsea Finn · Stanford · 2023
Framework
PyTorch
License
MIT
Stars
1,970
Summary
ACT employs a transformer-based action chunking policy and temporal ensembling to perform precise bimanual manipulation on a $5k ALOHA robot setup.
Abstract Summary
Key Points
- Transformer-based action chunking: predict a sequence of actions at once rather than one step at a time.
- Temporal ensembling averages overlapping actions across prediction windows to reduce jitter.
- Works with multi-view RGB + standard 14-DoF dual-arm robots; no depth or force-torque needed.
- Paired with ALOHA hardware: a $5k dual-arm teleoperation platform with full BoM released.
- Requires only 20–80 human demonstrations to learn new bimanual dexterous tasks.
Related Papers
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
Christen Millerdurai, Shaoxiang Wang, Yaxu Xie et al. · SIGGRAPH 2026 · May 2026
Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where sensing must...
MonoDuo: Using One Robot Arm to Learn Bimanual Policies
Sandeep Bajamahal, Lawrence Yunliang Chen, Toru Lin et al. · ICRA 2026 · May 2026
Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-arm robots, however,
SI-Diff: A Framework for Learning Search and High-Precision Insertion with a Force-Domain Diffusion Policy
Yibo Liu, Stanko Oparnica, Simon Shewchun-Jakaitis et al. · arXiv · May 2026
Contact-rich manipulation is fundamental in robotics but poses significant challenges due to uncertainties in relative poses, such as misalignments and small clearances in peg-in-hole tasks. Existing approaches typically address search and...
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
Matthew M. Hong, Jesse Zhang, Anusha Nagabandi et al. · arXiv · May 2026
Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action distributions that lack the coverage necessary for...