A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies
Mathilde Kappel, Clémence Grislain, Mohamed Chetouani, Olivier Sigaud, Louis Annabi, Faïz Ben Amar, Stéphane Doncieux, Mahdi Khoramshahi · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as ...
Abstract Summary
Key Points
- Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal input...
- Most state-of-the-art models predict actions in the end-effector pose space as sequences of actio...
- Training and evaluating these models requires large-scale collections of real-world demonstration...
- Collecting such data on real hardware typically relies on human teleoperation, making the process...
- We present an open-source sim-to-real experimental protocol that addresses this bottleneck: exper...
Method
Section III-A describes the paired simulated and real-world setup. Section III-B introduces the action chunk representation targeted by the pipeline, and how these chunks are processed and executed in the real robot. Section III-C details the expert and inference deployment protocols. Finally, Section III-D presents the underlying ROS2 interface.
In this work, we consider a paired simulated and real-world setup, composed of a Franka FR3 arm and a set of objects for which a URDF model is available, in our case, two colored cubes.
Experiments
Expert real robot data collection. We collected real-world data for VLA training with an action chunk size $k=8$ and $m=8$ chunks per trajectory i.e. $64$ actions per expert trajectory. We considered four language-conditioned combinations of two tasks ( push right , push left ) and two target objects (a red cube and a blue cube ).
We generate and deploy $50$ expert trajectories per task and compute metrics over the entire multi-task dataset composed of $200$ expert trajectories.
Sources and Demonstrations
Figures are reproduced from Kappel et al. See the full paper for experimental details and the project page, when the paper links one, for demonstrations and videos.
Related Papers
Capability-Aware Arbitration for Semantic Intent-Based Shared Control
Zhaoda Du, Michael Bowman, Xiaoli Zhang
Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (...
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al.
Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...
SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation
He Zhu, Lusen Zhao, Kwan Man Cheng et al.
Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines oft...
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Zhe Li, Zhenzhe Zhang, Yangyang Wei et al.
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...