Papers

Sorted by year (newest first)
DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog

foundation-models perception manipulation representation-learning vla
PDF Intermediate
No code repo Code updated: May 2026
Uncertainty Quantification for Flow-Based Vision-Language-Action Models

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

Ralf Römer, Maximilian Seeliger, Saida Liu et al. · RSS 2026 — Best Paper Award · Jun 2026

Quantifies epistemic uncertainty in flow-matching VLAs using velocity-field disagreement (VFD) across a small ensemble, enabling failure detection at deployment and sample-efficient active fine-tuning (SAVE).

vla foundation-models safety manipulation
Code PDF Advanced
GitHub ★ — Jun 2026
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction

Kai Xiong, Hongjie Fang, Lixin Yang et al. · arXiv · May 2026

Effectively handling the interplay between spatial perception and action generation remains a critical bottleneck in robotic manipulation. Existing methods typically treat spatial perception and action execution as decoupled or strictly...

imitation-learning manipulation foundation-models
PDF Intermediate
No code repo May 2026
DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset

DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset

Alexander Khazatsky, Karl Pertsch, Suraj Nair et al. · RSS 2024 · Jul 2024

A 350-hour dataset of diverse real-world robot manipulation across 22 robots in 71 scenes, designed to train scalable and generalist imitation learning policies.

imitation-learning manipulation foundation-models
Code PDF Advanced
GitHub ★ 362 Code updated: May 2026
Eureka: Human-Level Reward Design via Coding Large Language Models

Eureka: Human-Level Reward Design via Coding Large Language Models

Yecheng Jason Ma, William Liang, Guanzhi Wang et al. · ICLR 2024 (Oral) · May 2024

Eureka uses GPT-4 to write reward functions for RL environments, achieving human-level reward design on 29 tasks and enabling zero-shot sim-to-real transfer on Shadow Hand dexterous manipulation.

rl foundation-models sim-to-real
Code PDF Intermediate
GitHub ★ 3,161 Code updated: May 2026
LeRobot: A Library for Real-World Robot Learning

LeRobot: A Library for Real-World Robot Learning

Remi Cadene, Simon Alibert, Alexander Soare et al. · NeurIPS 2024 Workshop · Jun 2024

Hugging Face's LeRobot is an open-source PyTorch framework providing pretrained models, datasets, and training scripts for imitation and reinforcement learning on real robots, lowering the entry barrier to robot learning.

imitation-learning rl foundation-models manipulation
Code PDF Beginner
GitHub ★ 24,333 Code updated: May 2026
Octo: An Open-Source Generalist Robot Policy

Octo: An Open-Source Generalist Robot Policy

Dibya Ghosh, Homer Walke, Karl Pertsch et al. · RSS 2024 · Jul 2024

Octo is a large open-source transformer-based generalist robot policy trained on 800k trajectories, supporting language-conditioned and goal-image-conditioned control across 9 robotic platforms with efficient fine-tuning on consumer GPUs.

foundation-models manipulation imitation-learning open-source
Code PDF Intermediate
GitHub ★ 1,652 Code updated: May 2026
OK-Robot: Open-Ended Object Manipulation with Pretrained Vision-Language Models

OK-Robot: Open-Ended Object Manipulation with Pretrained Vision-Language Models

Peiqi Liu, Yat Long Lo, Ted Xiao et al. · arXiv preprint · Mar 2024

OK-Robot uses off-the-shelf VLMs (CLIP, OWL-ViT) and LLMs (GPT-4) to perform open-ended object manipulation in unseen homes without any training, achieving 58% success on real-world pick-and-place tasks.

foundation-models llm manipulation mobile-manipulation
Code PDF Intermediate
GitHub ★ 597 Code updated: May 2026
π0: A Vision-Language-Action Flow Model for General Robot Control

π0: A Vision-Language-Action Flow Model for General Robot Control

Karl Pertsch, Oliver Groth, Jonas Frey et al. · arXiv preprint · Oct 2024

π0 is a 3.5B-parameter VLA flow model from Physical Intelligence that achieves state-of-the-art general robot manipulation by mixing online RL with high-quality human demonstrations, available as an open-source PyTorch implementation.

foundation-models vla manipulation il
Code PDF Advanced
GitHub ★ 11,990 Code updated: May 2026
OpenVLA: An Open-Source Vision-Language-Action Model

OpenVLA: An Open-Source Vision-Language-Action Model

Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti et al. · arXiv · Jun 2024

A 7B-parameter open-source Vision-Language-Action model pre-trained on 970k real-world robot demonstrations, achieving strong generalization across robots and tasks.

vla llm-robotics foundation-models manipulation
Code PDF Intermediate
GitHub ★ 6,260 Code updated: May 2026
RT-2: Vision-Language-Action Models That Generalize to Novel Tasks

RT-2: Vision-Language-Action Models That Generalize to Novel Tasks

Anthony Brohan, Noah Brown, Justice Carbajal et al. · ICRA 2024 · May 2024

Google DeepMind's VLA model combining a vision-language foundation model with robot action outputs, showing emergent generalization to novel objects, backgrounds, and semantic instructions far beyond training data.

llm-robotics foundation-models manipulation
Code Advanced
GitHub ★ 1,853 Code updated: May 2026
Gymnasium: A Standard Interface for Reinforcement Learning Environments

Gymnasium: A Standard Interface for Reinforcement Learning Environments

Farama Foundation, Jordan Terry, Mark Towers et al. · JMLR 2023 · Mar 2023

Gymnasium is the maintained successor to OpenAI Gym, providing a standardized API for RL environments with 100+ built-in tasks, vectorized parallel execution, and native support for physics engines like MuJoCo, PyBullet, and IsaacGym.

rl foundation-models sim-to-real
Code Beginner
GitHub ★ 11,943 Code updated: May 2026
ManiSkill: A Unified Benchmark for Generalizable Manipulation Skills

ManiSkill: A Unified Benchmark for Generalizable Manipulation Skills

Jiayuan Gu, Sean Xiang, Stone Tao et al. · ICLR 2023 (Oral) · Jan 2023

ManiSkill is a GPU-parallelized robotics simulation benchmark with 20+ manipulation tasks, PartNet-Mobility assets, and unified observation/action spaces, supporting RL, IL, and VLA training at millions of steps per hour.

sim-to-real rl manipulation foundation-models
Code PDF Beginner
GitHub ★ 2,914 Code updated: May 2026
Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Open X-Embodiment Collaboration · arXiv · Oct 2023

The largest collaborative robot learning dataset ever assembled, spanning 22 robots from 21 institutions, enabling training of generalist RT-X policies that exhibit positive cross-embodiment transfer.

imitation-learning foundation-models manipulation dataset
Code PDF Intermediate
GitHub ★ 1,853 Code updated: May 2026
RT-1: Robotics Transformer for Real-World Control at Scale

RT-1: Robotics Transformer for Real-World Control at Scale

Anthony Brohan, Yevgen Chebotar, Chelsea Finn et al. · RSS 2023 · Jul 2023

RT-1 is a 35M-parameter transformer trained on 130K robot demonstrations that generalizes to new tasks, objects, and environments, forming the foundation for Google's RT-2 and RT-X line of VLA models.

foundation-models vla il manipulation
Code PDF Intermediate
GitHub ★ 1,723 Code updated: May 2026
RT-1: Robotics Transformer for Real-World Control at Scale

RT-1: Robotics Transformer for Real-World Control at Scale

Anthony Brohan, Noah Brown, Justice Carbajal et al. · RSS · Jul 2023

A large transformer model trained on 130k episodes of real robot manipulation to output discretized arm-and-gripper actions from RGB images and natural language instructions.

llm-robotics foundation-models imitation-learning
Code PDF Advanced
GitHub ★ 1,723 Code updated: May 2026
FAST-LIO2: Fast Direct LiDAR-Inertial Odometry

FAST-LIO2: Fast Direct LiDAR-Inertial Odometry

Wei Xu, Yixi Cai, Dongjiao He et al. · IEEE T-RO · 2022

A tightly-coupled LiDAR-inertial odometry system with incremental kd-tree mapping, enabling real-time state estimation for UAVs and mobile robots.

uav vla foundation-models
Code PDF Intermediate
GitHub ★ 4,707 Code updated: May 2026
CLIPort: What and Where Pathways for Robotic Manipulation

CLIPort: What and Where Pathways for Robotic Manipulation

Mohit Shridhar, Lucas Manuelli, Dieter Fox · CoRL 2021 · Nov 2021

CLIPort fuses CLIP's semantic understanding with Transporter Networks' spatial precision to perform language-conditioned manipulation tasks like 'stack the red block on the blue block' with 2D pick-and-place affordances.

foundation-models manipulation il
Code PDF Beginner
GitHub ★ 546 Code updated: May 2026

Suggested Learning Path

Read these papers in order to build expertise in Foundation Models.

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6

…and 27 more papers.