AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation
Mengfei Zhao, Dihong Huang, Yikai Tang, Peihao Li, Mingxuan Yan, Ruiqi Zhuang, Yanjia Huang, Jie Wang, Hai Zhai, Tony Zhou, Rui Zhang, Zhexi Luo, Yuchen Huang, Jianfei Yang, Jiachen Li · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine a...
Abstract Summary
Key Points
- Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet...
- We present AXIS, a growable community-driven data engine and benchmark for scalable robot learnin...
- The AXIS dataset currently contains 207 diverse tasks and 50K+ trajectories
- Meanwhile, AXIS organizes data into task snapshots and evaluates policies with a systematic held-...
- We compare vision-language-action (VLA) policies under a unified AXIS evaluation suite and analyz...
AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation
|Authors: Mengfei Zhao, Dihong Huang, Yikai Tang, Peihao Li, Mingxuan Yan, Ruiqi Zhuang, Yanjia Huang, Jie Wang, Hai Zhai, Tony Zhou, Rui Zhang, Zhexi Luo, Yuchen Huang, Jianfei Yang, Jiachen Li
|Venue: arXiv preprint | Year: 2026
|arXiv: 2607.21588v1
Abstract
Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine and benchmark for scalable robot learning, which enables browser-based teleoperation for large-scale demonstration collection, automatically generates and validates new manipulation tasks, and transforms community-collected demonstrations into training-ready data through automated success checking, quality filtering, trajectory smoothing, and visual and physics-based augmentation. The AXIS dataset currently contains 207 diverse tasks and 50K+ trajectories. Meanwhile, AXIS organizes data into task snapshots and evaluates policies with a systematic held-out protocol. We compare vision-language-action (VLA) policies under a unified AXIS evaluation suite and analyze scaling behavior across different data volumes. Continual pretraining on AXIS substantially improves the overall success rate of $π_{0.5}$ by 5.8%, outperforms the model pretrained on RoboCasa365 by 37.3%, and exhibits consistent scaling with increasing data volume, with the largest gains observed under layout, sensor-noise, and camera perturbations.
Key Contributions
- Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet…
- We present AXIS, a growable community-driven data engine and benchmark for scalable robot learnin…
- The AXIS dataset currently contains 207 diverse tasks and 50K+ trajectories
- Meanwhile, AXIS organizes data into task snapshots and evaluates policies with a systematic held-…
- We compare vision-language-action (VLA) policies under a unified AXIS evaluation suite and analyz…
Topics
- manipulation
- vision
- vla
- learning-from-demonstration
- benchmark
Code & Data
No code repository linked in paper metadata.
BibTeX
@article{Zhao2026_260721588v1,
title = {AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation},
author = {Mengfei Zhao and Dihong Huang and Yikai Tang and Peihao Li and Mingxuan Yan and Ruiqi Zhuang and Yanjia Huang and Jie Wang and Hai Zhai and Tony Zhou and Rui Zhang and Zhexi Luo and Yuchen Huang and Jianfei Yang and Jiachen Li},
year = {2026},
eprint = {2607.21588v1},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2607.21588v1}
}
Related Papers
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al. · arXiv preprint · Jul 2026
Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Zhe Li, Zhenzhe Zhang, Yangyang Wei et al. · arXiv preprint · Aug 2026
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
NVIDIA, :, Johan Bjorck et al. · arXiv preprint · Mar 2025
General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for...
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
Haoran Jiang, Jin Chen, Qingwen Bu et al. · arXiv preprint · Dec 2025
Humanoid robots require precise locomotion and dexterous manipulation to perform challenging loco-manipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This confines the robot to a limited workspace, preventing it from performing large...