GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
Chenghao Gu, Hanyang Yu, Jingbo Zhang, Haitao Lin, Wenyao Zhang, Jinghe Wang, Hanglei Jin, Shuzhao Xie, Jingyan Jiang, Zhi Wang · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse real-world environments remains costly and challenging. Action-conditioned world models offer a promising alternative, b...
Abstract Summary
Key Points
- Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen...
- Scaling robot learning and evaluation in diverse real-world environments remains costly and chall...
- Action-conditioned world models offer a promising alternative, but they often suffer from limited...
- To this end, we present GeniWorld, an interactive world model for robots that generalizes robustl...
- Building on pretrained video generative models, we use URDF-based rendering to transform numerica...
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
|Authors: Chenghao Gu, Hanyang Yu, Jingbo Zhang, Haitao Lin, Wenyao Zhang, Jinghe Wang, Hanglei Jin, Shuzhao Xie, Jingyan Jiang, Zhi Wang
|Venue: arXiv preprint | Year: 2026
|arXiv: 2608.06332v1
Abstract
Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse real-world environments remains costly and challenging. Action-conditioned world models offer a promising alternative, but they often suffer from limited action controllability and poor generalization to out-of-distribution (OOD) scenarios. To this end, we present GeniWorld, an interactive world model for robots that generalizes robustly across unseen scenarios. Building on pretrained video generative models, we use URDF-based rendering to transform numerical actions into visual action representations, enabling spatially grounded action control. By explicitly decoupling embodiment kinematics from environmental dynamics, our model mitigates scene overfitting and facilitates modeling of robot-environment interactions. To achieve closed-loop control, we construct an autoregressive video prediction model integrated with high-frequency robot kinematic control, enabling interaction with both robot policies and human teleoperators. In our experiments, even when trained solely on limited fixed-scene data, our model achieves superior in-domain performance and robust zero-shot generalization to highly randomized, unseen environments. For downstream applications, GeniWorld serves as a scalable policy evaluator that remains reliable under environmental perturbations. Furthermore, even with limited real-world demonstrations, GeniWorld generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in complex environments.
Key Contributions
- Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen…
- Scaling robot learning and evaluation in diverse real-world environments remains costly and chall…
- Action-conditioned world models offer a promising alternative, but they often suffer from limited…
- To this end, we present GeniWorld, an interactive world model for robots that generalizes robustl…
- Building on pretrained video generative models, we use URDF-based rendering to transform numerica…
Topics
- manipulation
- vision
- reinforcement-learning
- control
- learning-from-demonstration
- benchmark
Code & Data
No code repository linked in paper metadata.
BibTeX
@article{Gu2026_260806332v1,
title = {GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions},
author = {Chenghao Gu and Hanyang Yu and Jingbo Zhang and Haitao Lin and Wenyao Zhang and Jinghe Wang and Hanglei Jin and Shuzhao Xie and Jingyan Jiang and Zhi Wang},
year = {2026},
eprint = {2608.06332v1},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2608.06332v1}
}
Related Papers
Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators
Juan José García Cárdenas, Alperen Kenan, Hamidreza Raei et al. · arXiv preprint · Aug 2026
Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in challenging environments. In the nuclear industry, surface contact tasks such as swab sampling require precise path and force tracking, obstacle avoidance, and sustained operator attenti...
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al. · arXiv preprint · Jul 2026
Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Zhe Li, Zhenzhe Zhang, Yangyang Wei et al. · arXiv preprint · Aug 2026
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
NVIDIA, :, Johan Bjorck et al. · arXiv preprint · Mar 2025
General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for...