A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

Mathilde Kappel, Clémence Grislain, Mohamed Chetouani, Olivier Sigaud, Louis Annabi, Faïz Ben Amar, Stéphane Doncieux, Mahdi Khoramshahi · N/A · 2026

Framework

N/A

License

N/A

Stars

N/A

Summary

Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as ...

Abstract Summary

Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as sequences of action chunks. Training and evaluating these models requires large-scale collections of real-world demonstrations, pairing robot actions with the corresponding visual and proprioceptive observations. Collecting such data on real hardware typically relies on human teleoperation, making the process costly, time-consuming, and difficult to scale. We present an open-source sim-to-real experimental protocol that addresses this bottleneck: expert trajectories generated in simulation are replayed open-loop on a real Franka FR3 setup, where the corresponding real visual and proprioceptive observations are recorded and converted into a format compatible with VLA training. The same deployment stack is then reused, in closed-loop, to evaluate a trained policy on that setup, so that data collection and evaluation share an identical hardware configuration. Because each real recording is paired with the simulated trajectory that produced it, the protocol also yields a direct measurement of the sim-to-real gap. We release the collected datasets on Hugging Face together with the pipeline source code https://gitlab.isir.upmc.fr/kappel/sim2real_public_chunk_control.

Key Points

  • Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal input...
  • Most state-of-the-art models predict actions in the end-effector pose space as sequences of actio...
  • Training and evaluating these models requires large-scale collections of real-world demonstration...
  • Collecting such data on real hardware typically relies on human teleoperation, making the process...
  • We present an open-source sim-to-real experimental protocol that addresses this bottleneck: exper...

Method

Section III-A describes the paired simulated and real-world setup. Section III-B introduces the action chunk representation targeted by the pipeline, and how these chunks are processed and executed in the real robot. Section III-C details the expert and inference deployment protocols. Finally, Section III-D presents the underlying ROS2 interface.

In this work, we consider a paired simulated and real-world setup, composed of a Franka FR3 arm and a set of objects for which a URDF model is available, in our case, two colored cubes.

Method — A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies
Fig. 2 : Sim-to-real architecture ROS 2 closed-loop pipeline with Inference Server ( orange ) and Real-Time Controller Server ( blue ) for real-robot deployment.

Experiments

Expert real robot data collection. We collected real-world data for VLA training with an action chunk size $k=8$ and $m=8$ chunks per trajectory i.e. $64$ actions per expert trajectory. We considered four language-conditioned combinations of two tasks ( push right , push left ) and two target objects (a red cube and a blue cube ).

We generate and deploy $50$ expert trajectories per task and compute metrics over the entire multi-task dataset composed of $200$ expert trajectories.

Sources and Demonstrations

Figures are reproduced from Kappel et al. See the full paper for experimental details and the project page, when the paper links one, for demonstrations and videos.

Share

Related Papers

Capability-Aware Arbitration for Semantic Intent-Based Shared Control
arXiv preprint

Capability-Aware Arbitration for Semantic Intent-Based Shared Control

Zhaoda Du, Michael Bowman, Xiaoli Zhang

Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
Code PDF Intermediate
GitHub ★ —
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
arXiv preprint

DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al.

Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
PDF Advanced
SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation
arXiv preprint

SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

He Zhu, Lusen Zhao, Kwan Man Cheng et al.

Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines oft...

manipulation vla reinforcement-learning sim-to-real control learning-from-demonstration benchmark
PDF Advanced
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
arXiv preprint

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Zhe Li, Zhenzhe Zhang, Yangyang Wei et al.

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...

manipulation locomotion vision vla reinforcement-learning control learning-from-demonstration benchmark
PDF Advanced