Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation
FeaturedLuca Zanatta, Grzegorz Malczyk, Kostas Alexis · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
World models, learned generative models that predict how an environment evolves, have become a promising tool for sample-efficient robot learning. Yet how robust they are to environmental variability remains poorly understood. To address this, we conduct a systematic study using vision-based quadrot...
Abstract Summary
Key Points
- Uses simulation or synthetic data for training
- Builds predictive world models for planning
- Focuses on generalization across environments
- Incorporates tactile or force feedback for robust interaction
Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation
Authors: Luca Zanatta, Grzegorz Malczyk, Kostas Alexis
Venue: arXiv preprint | Year: 2026
arXiv: 2606.05015v1
Abstract
World models, learned generative models that predict how an environment evolves, have become a promising tool for sample-efficient robot learning. Yet how robust they are to environmental variability remains poorly understood. To address this, we conduct a systematic study using vision-based quadrotor navigation as a testbed problem, training DreamerV3-based world models under varying levels of environmental randomness and evaluating them across all levels through cross-environment validation, spanning both Self-Supervised Learning (SSL) pretraining and Reinforcement Learning (RL) fine-tuning. We then deploy all world models and associated navigation policies on a real quadrotor in unseen environments, including an open-loop run where the model receives just 2.5s of real sensory input before all sensors are cut off, leaving the system to navigate entirely in imagination over a 12m traverse. Our results show that world model robustness during SSL pretraining is a strong predictor of sim-to-real transfer: every model that generalized well in cross-environment SSL validation deployed successfully in the real world, passing through gaps as narrow as 0.67m, whereas the model that dominated simulation policy evaluation failed on the real platform. We further identify (a) the discrete latent size and (b) the training-sequence length as the dominant factors governing world model quality.
Key Contributions
- Uses simulation or synthetic data for training
- Builds predictive world models for planning
- Focuses on generalization across environments
- Incorporates tactile or force feedback for robust interaction
Topics
- sim-to-real
- reinforcement-learning
- navigation
- vision
- world-models
Code & Data
Code Repository: [https://github.com/ntnu-arl/world-model-nav-generalization. 2 Figure 2: Method overview. Fir](https://github.com/ntnu-arl/world-model-nav-generalization. 2 Figure 2: Method overview. Fir)
BibTeX
@article{Zanatta2026_260605015v1,
title={Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation},
author={Luca Zanatta and Grzegorz Malczyk and Kostas Alexis},
year={2026},
eprint={2606.05015v1},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2606.05015v1}
}
Related Papers
3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training
Jiaxin Shi, Xidong Zhang, Fucai Zhu et al. · arXiv preprint · Jun 2026
We propose a 3D-thinking-guided co-training framework that enables vision-language-action (VLA) models to perform 3D spatial reasoning implicitly during action prediction. Our core insight is that 3D geometry perception and 3D spatial reasoning are distinct capabilities that can be disentangled and ...
A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle
Gustavo Claudio Karl Couto, Eric Aislan Antonelo, Gabriel George Zipperer · arXiv preprint · Sep 2026
This paper presents a low-cost, open experimental platform for research in end-to-end autonomous driving with miniature Ackermann vehicles. The platform combines a physical vehicle, a printed urban track, data collection tools, trajectory registration, and a Webots digital twin, enabling controll...
GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
Tianyi Xie, Haotian Zhang, Jinhyung Park et al. · arXiv preprint · Jun 2026
Scaling humanoid loco-manipulation requires robot-compatible demonstrations across diverse objects, whole-body motions, and scene geometries, but teleoperation and motion capture are difficult to scale because each collection depends on physical setups, instrumented actors, and robot operation. We p...
LadderMan: Learning Humanoid Perceptive Ladder Climbing
Siheng Zhao, Yuanhang Zhang, Ziqi Lu et al. · arXiv preprint · Jun 2026
Humanoid robots hold great promise for operating in human-centered environments, yet ladder climbing remains one of the most challenging tasks due to sparse footholds and handholds, complex whole-body coordination, and sensitivity to perception and control errors. We present extbf{LadderMan}, a un...