AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
FeaturedWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li, Hang Yin, Huangxing Chen, Wenzhao Zheng, Jianjiang Feng, Jie Zhou, Jiwen Lu · Tsinghua University · 2026
Framework
N/A
License
N/A
Stars
29
Summary
AwareVLN introduces selective self-aware reasoning for VLN agents, triggering explicit spatial and progress analysis only at uncertain waypoints to improve robustness and explainability.
Abstract Summary
Key Points
- Selective self-aware reasoning for Vision-and-Language Navigation.
- Only triggers explicit reasoning at uncertain waypoints, saving computation.
- Maintains internal belief over task progress and spatial alignment.
- Improves success rates on R2R and REVERIE benchmarks.
- Produces human-interpretable reasoning traces for debugging.
Additional Notes
Overview
- Paper (PDF): 2605.22816
- arXiv: 2605.22816
Related Papers
- What Limits Vision-and-Language Navigation?
- Guide, Think and Act: Interactive Embodied Reasoning in VLA
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Related Papers
ALAM: Algebraically Consistent Latent Transitions for Vision-Language-Action Models
Zuojin Tang, Haoyun Liu, Xinyuan Chang et al. · arXiv · May 2026
Vision-Language-Action (VLA) models map visual observations and language instructions to robot actions through learned l...
DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
Xiaosong Jia, Bowen Yang, Zuhao Ge et al. · RSS 2026 · May 2026
Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLMs). Existing VLAs rely on end-to-end supervision to implicitly enable the action decoding process to learn...
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
Yajie Li, Bozhou Zhang, Chun Gu et al. · ICML 2026 · May 2026
Video foundation-models models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined futures for action execution remains challenging. Existing approaches...