Papers

Sorted by year (newest first)
3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training

Jiaxin Shi, Xidong Zhang, Fucai Zhu et al. · arXiv preprint · Jun 2026

We propose a 3D-thinking-guided co-training framework that enables vision-language-action (VLA) models to perform 3D spatial reasoning implicitly during action prediction. Our core insight is that 3D geometry perception and 3D spatial reasoning are distinct capabilities that can be disentangled and ...

sim-to-real reinforcement-learning vision vla manipulation
PDF Advanced
No code repo Jun 2026
Accelerating and Scaling MPC-Guided Reinforcement Learning for Humanoid Locomotion and Manipulation

Accelerating and Scaling MPC-Guided Reinforcement Learning for Humanoid Locomotion and Manipulation

Junheng Li, Liang Wu, Sergio A. Esteban et al. · arXiv preprint · Jun 2026

In humanoid motion control, model predictive control (MPC) offers physically grounded prediction and constraint handling, while reinforcement learning (RL) enables robust whole-body skills through large-scale simulation. However, using MPC inside RL often requires time-consuming problem construction or excessive training overhead, making such frameworks difficult to justify in practice. This work studies efficient training-time MPC guidance for humanoid locomotion and manipulation, termed MPC-RL. We introduce a centroidal-dynamics MPC reward formulation that leverages guidance from MPC trajectories in training time. To make this practical in massively parallel RL, we develop π^nMPC, a parallel-in-horizon and construction-free batched GPU MPC solver that operates directly on time-varying dynamics to avoid high memory usage and pre-compilation. Through a variety of comparative studies and hardware validations, we have found that MPC-RL achieves superior performance in locomotion and manipulation skills.

humanoid reinforcement-learning model-predictive-control locomotion manipulation
Code PDF Advanced
GitHub ★ — Jun 2026
Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

Sixu Yan, Shikang Wang, Binhua Huang et al. · arXiv preprint · Sep 2026

This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient genera...

manipulation vision reinforcement-learning
Code PDF Advanced
GitHub ★ — Sep 2026
AHEAD: Anticipatory Hand-Driven Teleoperation via Human Intent Prediction

AHEAD: Anticipatory Hand-Driven Teleoperation via Human Intent Prediction

Seok Joon Kim, Junho Lee, Federica Spinola et al. · arXiv preprint · Jul 2026

Direct hand-driven teleoperation maps an operator's hand motion to robot end-effector commands at every frame, enabling precise control, but it requires constant monitoring and correction during approach, grasp, and placement, which can be slow and fatiguing. For repetitive pick-and-place tasks, ...

manipulation reinforcement-learning planning control learning-from-demonstration
PDF Intermediate
No code repo Jul 2026
AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

Jianli Sun, Bin Tian, Qiyao Zhang et al. · arXiv preprint · Jun 2026

Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-level Unmanned Aerial Vehicle (UAV) movement and end-effector-level arm manipulation differ substantially in action scale, dynamics, and control objectives. In this paper, we propose AIR-VL...

uav vla manipulation
PDF Advanced
No code repo Jun 2026
AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

Mengfei Zhao, Dihong Huang, Yikai Tang et al. · arXiv preprint · Jul 2026

Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine a...

manipulation vision vla learning-from-demonstration benchmark
PDF Advanced
No code repo Jul 2026
ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models

ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models

Gehan Zheng, Matthew Johnson-Roberson, Weiming Zhi · arXiv preprint · Aug 2026

Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-camera setups: close gripper--object views help observe contact, but a poor approach may already push, miss, slip, or disturb the object before conventional de...

manipulation vision reinforcement-learning
PDF Intermediate
No code repo Aug 2026
ContactMimic: Humanoid Object Interaction via Contact Control

ContactMimic: Humanoid Object Interaction via Contact Control

Xinyao Li, Xialin He, Runpei Dong et al. · arXiv preprint · Jul 2026

Keypoint tracking alone is insufficient for object interaction tasks such as sitting on a chair, wiping a board, or pushing furniture, where the robot can reach the correct pose without making meaningful physical contact with the object. We present CONTACTMIMIC, a learning framework that tracks e...

manipulation reinforcement-learning control
PDF Intermediate
No code repo Jul 2026
Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation

Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation

Mehmet Turan Yardımcı · arXiv preprint · Jun 2026

Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy. A natural design choice is whether to use a single (unified) critic that estimates the combined value of all objectives, or separate (dual) critics with disjoint reward sign...

humanoid reinforcement-learning manipulation
PDF Intermediate
No code repo Jun 2026
Decoding Task Progress from VLA Representations

Decoding Task Progress from VLA Representations

Atiksh Bhardwaj, Edward Weiyi Duan, Prithwish Dan et al. · arXiv preprint · Aug 2026

Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools for understanding what these models represent internally or for monitoring them at runtime. Leveraging ideas from mechanistic interpretability, we...

manipulation vision vla reinforcement-learning control
PDF Intermediate
No code repo Aug 2026
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

Siyuan Ma, Boshi Zhang, Yutian Zhang et al. · arXiv preprint · Aug 2026

Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOW...

manipulation locomotion vision reinforcement-learning control benchmark
PDF Intermediate
No code repo Aug 2026
Deformable Object Manipulation under Partial Observability via Real-Time Full-Shape Estimation

Deformable Object Manipulation under Partial Observability via Real-Time Full-Shape Estimation

Kosar Behnia, Ville Kyrki, Gokhan Alcan · arXiv preprint · Sep 2026

Manipulating deformable objects (DOs) is challenging due to their high-dimensional state space, underactuated dynamics, and partial observability. In this paper, we propose cRVAE, a lightweight conditional recurrent variational autoencoder that estimates the full DO state from only partial corner...

manipulation planning control
PDF Intermediate
No code repo Sep 2026
Deliberate Practice: Learning Robot Skills under a Budget

Deliberate Practice: Learning Robot Skills under a Budget

Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut et al. · arXiv preprint · Aug 2026

We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorithm, \emph{Deliberate Practice (DP)}, that computes a provably \emph{budget-optimal} allocation---practicing skills that maximize expected ...

manipulation reinforcement-learning planning
PDF Intermediate
No code repo Aug 2026
Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators

Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators

Juan José García Cárdenas, Alperen Kenan, Hamidreza Raei et al. · arXiv preprint · Aug 2026

Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in challenging environments. In the nuclear industry, surface contact tasks such as swab sampling require precise path and force tracking, obstacle avoidance, and sustained operator attenti...

manipulation vision reinforcement-learning control learning-from-demonstration tactile benchmark
PDF Intermediate
No code repo Aug 2026
Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction

Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction

Thang Tran Viet, Thanh Nguyen Canh, Huy Uong Gia et al. · arXiv preprint · Sep 2026

Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in real-world environments. This paper introduces a humanoid robot prototype designed as a flexible testbed for developing and integrating artificial intelligence (AI) modules in HRI tasks. The...

manipulation vision vla reinforcement-learning control human-robot-interaction
Code PDF Intermediate
GitHub ★ — Sep 2026
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al. · arXiv preprint · Jul 2026

Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
PDF Advanced
No code repo Jul 2026
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

DreamX Team, Rui Chen, Xiangxiang Chu et al. · arXiv preprint · Aug 2026

We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism al...

manipulation reinforcement-learning control
PDF Intermediate
No code repo Aug 2026
Dual Advantage Fields

Dual Advantage Fields

Alexey Zemtsov, Maxim Bobrin, Alexander Nikulin et al. · ICML 2026 · Jun 2026

Offline goal-conditioned reinforcement learning requires both long-horizon reachability estimates and local action comparisons. Dual goal representations provide value fields that capture global goal reachability, but they do not directly specify which action should be preferred at a given state. We...

reinforcement-learning manipulation locomotion
PDF Advanced
No code repo Jun 2026
DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog

foundation-models perception manipulation representation-learning vla
PDF Intermediate
No code repo Code updated: May 2026
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

Junfeng Li, Junjie He, Zhide Zhong et al. · arXiv preprint · Aug 2026

Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse v...

manipulation vision vla reinforcement-learning control benchmark
PDF Intermediate
No code repo Aug 2026
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

Christen Millerdurai, Shaoxiang Wang, Yaxu Xie et al. · SIGGRAPH 2026 · May 2026

Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where sensing must...

manipulation manipulation simulation
Code PDF Intermediate
Code ★ 0 May 2026
Evidence-Gated Task and Motion Planning with Vision-Language Models

Evidence-Gated Task and Motion Planning with Vision-Language Models

Tsunehiko Tanaka, Matthew Stephenson, Alistair Macvicar et al. · arXiv preprint · Aug 2026

Robots executing long-horizon manipulation tasks from natural-language instructions must reason about both semantic task structure and geometric feasibility. However, under partial observability, the availability of goal-relevant objects may be uncertain. In such cases, approaches that combine Vi...

manipulation vision vla reinforcement-learning planning
PDF Intermediate
No code repo Aug 2026
FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation

FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation

Lifeng Zhuo, Wendi Chen, Han Xue et al. · arXiv preprint · Jul 2026

In contact-rich manipulation, action multimodality and reactivity dominate different stages of a single episode. Before contact, multiple trajectories might be equally valid, making it important to preserve diverse action modes. After contact, geometric constraints and force limits narrow the sol...

manipulation vision reinforcement-learning control
Code PDF Advanced
GitHub ★ — Jul 2026
FabriVLA: A Lightweight Vision-Language-Action Model for Precise Multi-Task Manipulation

FabriVLA: A Lightweight Vision-Language-Action Model for Precise Multi-Task Manipulation

Shiyuan Yang, Borong Zhang, Jizheng Zhang et al. · arXiv preprint · Jul 2026

We present FabriVLA, a lightweight Vision-Language-Action model for Precise Multi-Task Manipulation. FabriVLA combines an InternVL3.5 vision-language backbone with a flow-matching action head featuring gated self-attention across action tokens and shallow VLM layer fusion for enriched spatial con...

manipulation vision vla reinforcement-learning benchmark
PDF Intermediate
No code repo Jul 2026
FasTac: A Curved Multispectral Vision-Based Tactile Sensor for High-Speed High-Precision 3D Shape and Force Perception

FasTac: A Curved Multispectral Vision-Based Tactile Sensor for High-Speed High-Precision 3D Shape and Force Perception

Xiaofan Lu, Kaiji Huang, Jiahui Chen et al. · arXiv preprint · Jul 2026

Curved tactile fingertips for dexterous manipulation must resolve fine contact geometry, distinguish normal and tangential loads, and capture transient signals. Existing curved vision-based tactile sensors struggle to combine accurate 3D reconstruction, three-axis force estimation, and high-speed...

manipulation vision tactile
PDF Intermediate
No code repo Jul 2026
Flash-WAM: Modality-Aware Distillation for World Action Models

Flash-WAM: Modality-Aware Distillation for World Action Models

Arman Akbari, Ci Zhang, Arash Akbari et al. · arXiv preprint · Jun 2026

World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off...

sim-to-real reinforcement-learning diffusion-policy manipulation humanoid
PDF Intermediate
No code repo Jun 2026
FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference

FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference

Zekai Li, Jiaming Tang, Zhijian Liu · arXiv preprint · Aug 2026

Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable asynchronous execution. This challenge is particularly pronounced in flow-matching-based VLA models, where action de...

manipulation vision vla reinforcement-learning control
Code PDF Intermediate
GitHub ★ — Aug 2026
Aligning Flow Map Policies with Optimal Q-Guidance

Aligning Flow Map Policies with Optimal Q-Guidance

Christos Ziakas, Alessandra Russo, Avishek Joey Bose · arXiv · May 2026

Generative policies based on expressive model classes, such as diffusion-models and flow matching, are well-suited to complex control problems with highly multimodal action distributions. Their expressivity, however, comes at a significant inference cost:...

rl diffusion-models manipulation
PDF Advanced
No code repo May 2026
FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

Chenhuan Liu, Yi Xu, Feng Wu et al. · arXiv preprint · Sep 2026

Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-object manipulation, where policies must track changing s...

manipulation vision vla reinforcement-learning control benchmark
PDF Intermediate
No code repo Sep 2026
GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training

GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training

Zhihai Bi, Qiang Zhang, Guoyang Zhao et al. · arXiv preprint · Jun 2026

Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid tr...

humanoid manipulation vision
PDF Intermediate
No code repo Jun 2026
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

Chenghao Gu, Hanyang Yu, Jingbo Zhang et al. · arXiv preprint · Aug 2026

Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse real-world environments remains costly and challenging. Action-conditioned world models offer a promising alternative, b...

manipulation vision reinforcement-learning control learning-from-demonstration benchmark
PDF Intermediate
No code repo Aug 2026
GesVLA: Gesture-Aware Vision-Language-Action Model

GesVLA: Gesture-Aware Vision-Language-Action Model

Wenxuan Guo, Ziyuan Li, Meng Zhang et al. · arXiv preprint · May 2026

GesVLA augments standard VLA models with gesture awareness, enabling robots to interpret verbal instructions alongside human hand gestures for disambiguated manipulation.

vla manipulation hand-tracking
Code PDF Intermediate
GitHub ★ 18 Code updated: May 2026
GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

Yupeng Zheng, Xiang Li, Songen Gu et al. · arXiv preprint · Sep 2026

Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction objectives may omit critical physical and task structure while retaining control-irrelevant visual redundancy. We call th...

manipulation vision vla reinforcement-learning control
PDF Intermediate
No code repo Sep 2026
GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

Tianyi Xie, Haotian Zhang, Jinhyung Park et al. · arXiv preprint · Jun 2026

Scaling humanoid loco-manipulation requires robot-compatible demonstrations across diverse objects, whole-body motions, and scene geometries, but teleoperation and motion capture are difficult to scale because each collection depends on physical setups, instrumented actors, and robot operation. We p...

sim-to-real reinforcement-learning vision manipulation humanoid
Code PDF Advanced
GitHub ★ 275 Code updated: Jun 2026
Grasp, Handover, Rotate: Bimanual Object Reorientation via Compositional Diffusion and Energy-Based Optimization

Grasp, Handover, Rotate: Bimanual Object Reorientation via Compositional Diffusion and Energy-Based Optimization

Wun Lam Yeung, Wenjun Liu, Yui Cheung Yu et al. · arXiv preprint · Jul 2026

Bimanual object reorientation - picking an object, handing it over between two arms, and placing it in a desired target pose - is valuable when direct placement from the initial grasp is infeasible due to collisions, kinematic constraints, or poor final orientation. However, achieving this under ...

manipulation reinforcement-learning sim-to-real planning
PDF Advanced
No code repo Jul 2026
GS-Agent: Creating 4D Physical Worlds With Generative Simulation

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

Hongxin Zhang, Chunru Lin, Junyan Li et al. · arXiv preprint · Jul 2026

Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in gene...

manipulation vision reinforcement-learning control
Code PDF Advanced
GitHub ★ — Jul 2026
HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

Zengjue Chen, Peidong Liu, Jiawei Li et al. · arXiv preprint · Sep 2026

Generalist robot policies have demonstrated strong generalization across robotic manipulation tasks, yet their success rates remain limited in com- plex long-horizon scenarios. Recent methods improve Visual-Language-Action (VLA) policies through online reinforcement learning on real robots, but s...

manipulation vision vla reinforcement-learning benchmark
PDF Intermediate
No code repo Sep 2026
Improving Robotic Generalist Policies via Flow Reversal Steering

Improving Robotic Generalist Policies via Flow Reversal Steering

Andy Tang, William Chen, Andrew Wagenmaker et al. · arXiv preprint · Jun 2026

Generalist policies can learn a wide range of skills from diverse robot datasets. In order to solve or improve on challenging news tasks, we need a way to infer and invoke the appropriate actions from the policy's rich behavioral prior, especially when directly commanding the policy fails. We focus ...

manipulation vla world-models
PDF Advanced
No code repo Jun 2026
LadderMan: Learning Humanoid Perceptive Ladder Climbing

LadderMan: Learning Humanoid Perceptive Ladder Climbing

Siheng Zhao, Yuanhang Zhang, Ziqi Lu et al. · arXiv preprint · Jun 2026

Humanoid robots hold great promise for operating in human-centered environments, yet ladder climbing remains one of the most challenging tasks due to sparse footholds and handholds, complex whole-body coordination, and sensitivity to perception and control errors. We present extbf{LadderMan}, a un...

sim-to-real reinforcement-learning vision manipulation humanoid
PDF Advanced
No code repo Jun 2026
M3imic: Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking

M3imic: Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking

Zuxing Lu, Ziang Zheng, Yao Lyu et al. · arXiv preprint · Jun 2026

Building a general-purpose whole-body controller is essential for enabling diverse motion capabilities in humanoid robots across a wide range of downstream tasks, including locomotion and loco-manipulation. Different tasks rely on distinct motion reference modalities: locomotion primarily depends on...

sim-to-real reinforcement-learning locomotion manipulation humanoid
Code PDF Advanced
GitHub ★ — Jun 2026
Memory as Plans: World-Action Modeling with Memory-Grounded Planning

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

Sizhe Zhao, Haozhe Xie, Weiyu Zhao et al. · arXiv preprint · Sep 2026

Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their...

manipulation vision reinforcement-learning planning
PDF Intermediate
No code repo Sep 2026
MemoryWAM: Efficient World Action Modeling with Persistent Memory

MemoryWAM: Efficient World Action Modeling with Persistent Memory

Sizhe Yang, Juncheng Mu, Tianming Wei et al. · arXiv preprint · Jun 2026

Robust robotic manipulation in the real world requires not only an understanding of the current observation, but also memory and dynamics modeling. World action models (WAMs) posse...

ar-vr gpu-acceleration manipulation
PDF Intermediate
No code repo Jun 2026
MemoryWAM: Efficient World Action Modeling with Persistent Memory

MemoryWAM: Efficient World Action Modeling with Persistent Memory

Sizhe Yang, Juncheng Mu, Tianming Wei et al. · arXiv preprint · Jun 2026

Robust robotic manipulation in the real world requires not only an understanding of the current observation, but also memory and dynamics modeling. World action models (WAMs) posse...

ar-vr gpu-acceleration manipulation
PDF Intermediate
No code repo Jun 2026
MonoDuo: Using One Robot Arm to Learn Bimanual Policies

MonoDuo: Using One Robot Arm to Learn Bimanual Policies

Sandeep Bajamahal, Lawrence Yunliang Chen, Toru Lin et al. · ICRA 2026 · May 2026

Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-arm robots, however,

manipulation bimanual imitation-learning dataset
PDF Intermediate
No code repo Code updated: May 2026
Native Video-Action Pretraining for Generalizable Robot Control

Native Video-Action Pretraining for Generalizable Robot Control

Qihang Zhang, Lin Li, Luyao Zhang et al. · arXiv preprint · Jul 2026

The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments. To bridge this gap, we present LingBot-VA 2.0, a video-action foun...

manipulation vision reinforcement-learning control
PDF Advanced
No code repo Jul 2026
Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

Kinam Kim, Namiko Saito, Heecheol Kim et al. · arXiv preprint · Jun 2026

Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions du...

ar-vr manipulation rl
PDF Intermediate
No code repo Jun 2026
Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

Kinam Kim, Namiko Saito, Heecheol Kim et al. · arXiv preprint · Jun 2026

Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions du...

ar-vr manipulation rl
PDF Intermediate
No code repo Jun 2026
PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

Sergey Arkhangelskiy · arXiv preprint · May 2026

Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with N <= 25 rollouts per condition, almost always without confidence intervals or

benchmarking vla manipulation evaluation
PDF Intermediate
No code repo Code updated: May 2026
Quasi-static analysis of passive stability in a novel underactuated multi-finger hand

Quasi-static analysis of passive stability in a novel underactuated multi-finger hand

Léonie Plancoulaine, Sylvain Guégan, Franck Plestan et al. · arXiv preprint · Sep 2026

Underactuated robotic hands achieve adaptive and robust grasping with a reduced number of actuators, but predicting the stable equilibrium pose of the grasped object remains a significant challenge. This paper introduces a quasi-static analytical approach to assess passive stability in underactua...

manipulation
PDF Intermediate
No code repo Sep 2026
Rapid Learning of Dexterous In-Hand Pen Writing through Real-Time Jacobian Estimation

Rapid Learning of Dexterous In-Hand Pen Writing through Real-Time Jacobian Estimation

Kai Stewart, Yasunori Toshimitsu, Robert K. Katzschmann · arXiv preprint · Sep 2026

Dexterous in-hand manipulation of a grasped object with an anthropomorphic hand is an unsolved frontier for robot dexterity. The contact-richness and highly dynamic nature of object-hand interactions tend to require extensive modeling or data-collection efforts for learning-based approaches. Mode...

manipulation reinforcement-learning control learning-from-demonstration
Code PDF Intermediate
GitHub ★ — Sep 2026
Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

Yuxuan Chen, Wanruo Zhang, Xiao Li · arXiv preprint · Aug 2026

Vision-Language-Action (VLA) models have recently achieved promising performance in robotic manipulation. However, existing benchmarks mainly evaluate generalization on static manipulation tasks and largely overlook dynamic interaction scenarios. To address this gap, we present ReflexBench, a ben...

manipulation vision vla reinforcement-learning control benchmark
Code PDF Advanced
GitHub ★ — Aug 2026
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

Zhenxuan Fan, Bo Zhang, Yutong Lin et al. · arXiv preprint · Sep 2026

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedura...

manipulation vision vla planning benchmark
Code PDF Advanced
GitHub ★ — Sep 2026
RoboTTT: Context Scaling for Robot Policies

RoboTTT: Context Scaling for Robot Policies

Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng et al. · arXiv preprint · Jul 2026

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, witho...

manipulation vision vla reinforcement-learning learning-from-demonstration
Code PDF Advanced
GitHub ★ — Jul 2026
Safety-aware Skill Adaptation for Reinforcement Learning in Dynamic Environments

Safety-aware Skill Adaptation for Reinforcement Learning in Dynamic Environments

A K M Nadimul Haque, Sheila Sutjipto, Marc G. Carmichael et al. · arXiv preprint · Sep 2026

Skill adaptation frameworks based on reinforcement learning often require restrictive assumptions to maintain stability, such as fixed observations or tightly controlled exploration schedules. In cluttered and dynamic environments, however, unrestricted exploration can lead to unsafe behaviour an...

manipulation reinforcement-learning control
PDF Intermediate
No code repo Sep 2026
Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

Shutong Ding, Zejia Zhong, Zhongyi Wang et al. · ICML 2026 · May 2026

Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approaches, one representativ

diffusion-policy reinforcement-learning locomotion manipulation
PDF Advanced
No code repo Code updated: May 2026
Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

Yu Qi, Zhang Ye, Xinyi Xu et al. · arXiv preprint · Jul 2026

Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferring to salient cues rather than grounding language. We introduce a diagnostic framework that localizes this failure to individual \textit{instruction...

manipulation reinforcement-learning learning-from-demonstration benchmark
PDF Intermediate
No code repo Jul 2026
SEED-UMI: Sharing the Exoskeleton between human and robot for onE-to-one Dexterous demonstration

SEED-UMI: Sharing the Exoskeleton between human and robot for onE-to-one Dexterous demonstration

Tengbo Yu, Jiahao Wu, Daohan Li et al. · arXiv preprint · Sep 2026

Imitation learning for dexterous hands is bottlenecked by the difficulty of collecting contact-rich demonstrations that transfer faithfully to the robot. Prior wearable-exoskeleton systems record only on the human side and retarget via open-loop mappings calibrated in free space, which degrade un...

manipulation vision reinforcement-learning learning-from-demonstration
Code PDF Intermediate
GitHub ★ — Sep 2026
SI-Diff: A Framework for Learning Search and High-Precision Insertion with a Force-Domain Diffusion Policy

SI-Diff: A Framework for Learning Search and High-Precision Insertion with a Force-Domain Diffusion Policy

Yibo Liu, Stanko Oparnica, Simon Shewchun-Jakaitis et al. · arXiv · May 2026

Contact-rich manipulation is fundamental in robotics but poses significant challenges due to uncertainties in relative poses, such as misalignments and small clearances in peg-in-hole tasks. Existing approaches typically address search and...

manipulation tactile manipulation
Code PDF Intermediate
Code ★ 0 May 2026
Synthetic Data Generation and Vision-based Wrinkle and Keypoint Detection for Bimanual Cloth Manipulation

Synthetic Data Generation and Vision-based Wrinkle and Keypoint Detection for Bimanual Cloth Manipulation

Ariel Herrera, Xueyang Kang, Atal Anil Kumar · arXiv preprint · Jun 2026

Robotic manipulation of textiles remains challenging because continuous deformation and self-occlusions hinder the robust visual perception required to estimate the cloth's state. To address the lack of annotated real-world data, we developed a Blender-based synthetic pipeline exporting auto-annotat...

vision manipulation reinforcement-learning
PDF Intermediate
No code repo Jun 2026
Task-space model-based control of pneumatic soft actuators

Task-space model-based control of pneumatic soft actuators

Nithin S. Kumar, Joshua Gaston, D. Caleb Rucker et al. · arXiv preprint · Aug 2026

Soft actuators enable dexterous and compliant interaction, but closed-loop task-space control remains challenging due to strong nonlinearities, distributed deformation, and uncertainty in their dynamics. This paper presents a real-time dynamic-model-based task-space feedback and estimation framew...

manipulation control
PDF Advanced
No code repo Aug 2026
Tensegrity Continuum Robots Enable Task-Adaptive Morphologies for Cooperative Behaviors

Tensegrity Continuum Robots Enable Task-Adaptive Morphologies for Cooperative Behaviors

Mahmud Hasan Saikot, Sydney Spiegel, Sudheera Akalanka Kariyawasam et al. · arXiv preprint · Aug 2026

Robots that can change their morphologies and behaviors for different tasks and environments hold great promise for adaptable, multifunctional systems. Modular reconfigurable robots (MRRs) can achieve such functionalities by docking and rearranging individual units, but most rely on rigid modules...

manipulation locomotion reinforcement-learning
PDF Intermediate
No code repo Aug 2026
THRIVE: Therapeutic Humanoid Robot In Virtual Environment

THRIVE: Therapeutic Humanoid Robot In Virtual Environment

Jin Xu, Yu-Ping Chen, Ayanna Howard · arXiv preprint · Aug 2026

This paper presents THRIVE (Therapeutic Humanoid Robot In Virtual Environment), an at-home rehabilitation platform that integrates a suite of virtual-reality upper-body rehabilitation games, a real-time camera-based motion-tracking system, and a socially interactive robot therapist. The system is...

manipulation vision reinforcement-learning human-robot-interaction
PDF Intermediate
No code repo Aug 2026
Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

Vivek Chavan, Yahuan Shi, Oliver Heimann et al. · arXiv preprint · Sep 2026

Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA ...

manipulation vision vla reinforcement-learning control learning-from-demonstration
PDF Intermediate
No code repo Sep 2026
Uncertainty Quantification for Flow-Based Vision-Language-Action Models

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

Ralf Römer, Maximilian Seeliger, Saida Liu et al. · RSS 2026 — Best Paper Award · Jun 2026

Quantifies epistemic uncertainty in flow-matching VLAs using velocity-field disagreement (VFD) across a small ensemble, enabling failure detection at deployment and sample-efficient active fine-tuning (SAVE).

vla foundation-models safety manipulation
Code PDF Advanced
GitHub ★ — Jun 2026
UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition Modeling

UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition Modeling

Wei Li, Rui Shao, Jie He et al. · arXiv preprint · Sep 2026

Recent advances in Vision-Language-Action (VLA) models have improved robotic manipulation, yet observation-to-action learning remains limited by a fundamental transition realizability gap, manifested in three tightly coupled problems: (i) Transition ambiguity. Visually similar current observation...

manipulation vision vla reinforcement-learning
Code PDF Intermediate
GitHub ★ — Sep 2026
Video2DoorTraversal: Push Door Traversal via Simulated Door Twins

Video2DoorTraversal: Push Door Traversal via Simulated Door Twins

Xincheng Tang, Yiji Chen, Youhan Xie et al. · arXiv preprint · Aug 2026

Door opening and traversal is a long-horizon loco-manipulation task that requires precise handle interaction and coordinated base-arm control. We present Video2DoorTraversal, a single-video real-to-sim-to-real framework for wheel-legged mobile manipulators. Given one RGB video of a real door, Doo...

manipulation locomotion vision reinforcement-learning sim-to-real control learning-from-demonstration
PDF Intermediate
No code repo Aug 2026
VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations

VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations

Hisham Khalil, Neil Fernandes, Thomas M. Kwok et al. · arXiv preprint · Aug 2026

Contact-rich manipulation requires precise tracking and mechanical compliance, where variable impedance control can improve robustness in task success, whereas static compliance cannot adapt to varying contact constraints. Variable impedance skills can be learned from demonstrations, avoiding com...

manipulation reinforcement-learning control learning-from-demonstration
PDF Advanced
No code repo Aug 2026
VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies

VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies

Chang Song, Bin Qian, Yan Feng et al. · arXiv preprint · Sep 2026

Long-horizon robot manipulation with Vision-Language-Action (VLA) policies remains vulnerable to execution-time deviations, as final task success provides little information for diagnosing and correcting failures caused by action noise, object displacement, or goal misalignment. We introduce a st...

manipulation vision vla reinforcement-learning benchmark
PDF Intermediate
No code repo Sep 2026
What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

Vivek Chavan, Pengtao Xie, Yahuan Shi et al. · arXiv preprint · Sep 2026

Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes w...

manipulation vision vla reinforcement-learning control
PDF Intermediate
No code repo Sep 2026
WT-UMI: Tactile-based Whole-Body Manipulation via Force-Supervised Contact-Aware Planning

WT-UMI: Tactile-based Whole-Body Manipulation via Force-Supervised Contact-Aware Planning

Jaehwi Jang, Zhaoyuan Gu, Alfred Cueva et al. · arXiv preprint · Jun 2026

Whole-body humanoid manipulation of bulky, deformable, and shared-load objects requires distributed contact sensing and explicit force regulation, yet most imitation policies treat contact force only implicitly. On the other hand, different demonstration sources provide complementary modalities with...

humanoid manipulation tactile
PDF Advanced
No code repo Jun 2026
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction

Kai Xiong, Hongjie Fang, Lixin Yang et al. · arXiv · May 2026

Effectively handling the interplay between spatial perception and action generation remains a critical bottleneck in robotic manipulation. Existing methods typically treat spatial perception and action execution as decoupled or strictly...

imitation-learning manipulation foundation-models
PDF Intermediate
No code repo May 2026
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Zhe Li, Zhenzhe Zhang, Yangyang Wei et al. · arXiv preprint · Aug 2026

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...

manipulation locomotion vision vla reinforcement-learning control learning-from-demonstration benchmark
PDF Advanced
No code repo Aug 2026
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

NVIDIA, :, Johan Bjorck et al. · arXiv preprint · Mar 2025

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
Code PDF Advanced
GitHub ★ — Mar 2025
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control

WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control

Haoran Jiang, Jin Chen, Qingwen Bu et al. · arXiv preprint · Dec 2025

Humanoid robots require precise locomotion and dexterous manipulation to perform challenging loco-manipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This confines the robot to a limited workspace, preventing it from performing large...

manipulation locomotion vision vla reinforcement-learning control learning-from-demonstration benchmark
Code PDF Advanced
GitHub ★ — Dec 2025
DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset

DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset

Alexander Khazatsky, Karl Pertsch, Suraj Nair et al. · RSS 2024 · Jul 2024

A 350-hour dataset of diverse real-world robot manipulation across 22 robots in 71 scenes, designed to train scalable and generalist imitation learning policies.

imitation-learning manipulation foundation-models
Code PDF Advanced
GitHub ★ 362 Code updated: May 2026
LeRobot: A Library for Real-World Robot Learning

LeRobot: A Library for Real-World Robot Learning

Remi Cadene, Simon Alibert, Alexander Soare et al. · NeurIPS 2024 Workshop · Jun 2024

Hugging Face's LeRobot is an open-source PyTorch framework providing pretrained models, datasets, and training scripts for imitation and reinforcement learning on real robots, lowering the entry barrier to robot learning.

imitation-learning rl foundation-models manipulation
Code PDF Beginner
GitHub ★ 24,333 Code updated: May 2026
Octo: An Open-Source Generalist Robot Policy

Octo: An Open-Source Generalist Robot Policy

Dibya Ghosh, Homer Walke, Karl Pertsch et al. · RSS 2024 · Jul 2024

Octo is a large open-source transformer-based generalist robot policy trained on 800k trajectories, supporting language-conditioned and goal-image-conditioned control across 9 robotic platforms with efficient fine-tuning on consumer GPUs.

foundation-models manipulation imitation-learning open-source
Code PDF Intermediate
GitHub ★ 1,652 Code updated: May 2026
OK-Robot: Open-Ended Object Manipulation with Pretrained Vision-Language Models

OK-Robot: Open-Ended Object Manipulation with Pretrained Vision-Language Models

Peiqi Liu, Yat Long Lo, Ted Xiao et al. · arXiv preprint · Mar 2024

OK-Robot uses off-the-shelf VLMs (CLIP, OWL-ViT) and LLMs (GPT-4) to perform open-ended object manipulation in unseen homes without any training, achieving 58% success on real-world pick-and-place tasks.

foundation-models llm manipulation mobile-manipulation
Code PDF Intermediate
GitHub ★ 597 Code updated: May 2026
π0: A Vision-Language-Action Flow Model for General Robot Control

π0: A Vision-Language-Action Flow Model for General Robot Control

Karl Pertsch, Oliver Groth, Jonas Frey et al. · arXiv preprint · Oct 2024

π0 is a 3.5B-parameter VLA flow model from Physical Intelligence that achieves state-of-the-art general robot manipulation by mixing online RL with high-quality human demonstrations, available as an open-source PyTorch implementation.

foundation-models vla manipulation il
Code PDF Advanced
GitHub ★ 11,990 Code updated: May 2026
OpenVLA: An Open-Source Vision-Language-Action Model

OpenVLA: An Open-Source Vision-Language-Action Model

Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti et al. · arXiv · Jun 2024

A 7B-parameter open-source Vision-Language-Action model pre-trained on 970k real-world robot demonstrations, achieving strong generalization across robots and tasks.

vla llm-robotics foundation-models manipulation
Code PDF Intermediate
GitHub ★ 6,260 Code updated: May 2026
RT-2: Vision-Language-Action Models That Generalize to Novel Tasks

RT-2: Vision-Language-Action Models That Generalize to Novel Tasks

Anthony Brohan, Noah Brown, Justice Carbajal et al. · ICRA 2024 · May 2024

Google DeepMind's VLA model combining a vision-language foundation model with robot action outputs, showing emergent generalization to novel objects, backgrounds, and semantic instructions far beyond training data.

llm-robotics foundation-models manipulation
Code Advanced
GitHub ★ 1,853 Code updated: May 2026
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Cheng Chi, Siyuan Feng, Yilun Du et al. · RSS 2023 · Jul 2023

A behavior cloning approach that models robot policies as conditional diffusion processes, enabling multimodal action distributions and smooth action sequences.

imitation-learning diffusion-models manipulation behavior-cloning
Code PDF Intermediate
GitHub ★ 4,194 Code updated: May 2026
Stretch: A Small, Capable, and Affordable Mobile Manipulator

Stretch: A Small, Capable, and Affordable Mobile Manipulator

Charlie C. Kemp, Bilsen D. Kim, Blake S. Kim et al. · ICRA Workshop · 2023

A low-cost mobile manipulator designed for indoor home and office environments, with a large open-source ecosystem and teleoperation-based data collection.

manipulation uav imitation-learning
Code PDF Beginner
GitHub ★ 1,200 Code updated: Dec 2024
ManiSkill: A Unified Benchmark for Generalizable Manipulation Skills

ManiSkill: A Unified Benchmark for Generalizable Manipulation Skills

Jiayuan Gu, Sean Xiang, Stone Tao et al. · ICLR 2023 (Oral) · Jan 2023

ManiSkill is a GPU-parallelized robotics simulation benchmark with 20+ manipulation tasks, PartNet-Mobility assets, and unified observation/action spaces, supporting RL, IL, and VLA training at millions of steps per hour.

sim-to-real rl manipulation foundation-models
Code PDF Beginner
GitHub ★ 2,914 Code updated: May 2026
Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Open X-Embodiment Collaboration · arXiv · Oct 2023

The largest collaborative robot learning dataset ever assembled, spanning 22 robots from 21 institutions, enabling training of generalist RT-X policies that exhibit positive cross-embodiment transfer.

imitation-learning foundation-models manipulation dataset
Code PDF Intermediate
GitHub ★ 1,853 Code updated: May 2026
RT-1: Robotics Transformer for Real-World Control at Scale

RT-1: Robotics Transformer for Real-World Control at Scale

Anthony Brohan, Yevgen Chebotar, Chelsea Finn et al. · RSS 2023 · Jul 2023

RT-1 is a 35M-parameter transformer trained on 130K robot demonstrations that generalizes to new tasks, objects, and environments, forming the foundation for Google's RT-2 and RT-X line of VLA models.

foundation-models vla il manipulation
Code PDF Intermediate
GitHub ★ 1,723 Code updated: May 2026
CLIPort: What and Where Pathways for Robotic Manipulation

CLIPort: What and Where Pathways for Robotic Manipulation

Mohit Shridhar, Lucas Manuelli, Dieter Fox · CoRL 2021 · Nov 2021

CLIPort fuses CLIP's semantic understanding with Transporter Networks' spatial precision to perform language-conditioned manipulation tasks like 'stack the red block on the blue block' with 2D pick-and-place affordances.

foundation-models manipulation il
Code PDF Beginner
GitHub ★ 546 Code updated: May 2026