PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
Xinyu Guo, Bin Xie, Wei Chai, Xianchi Deng, Ruoxi Jia · · 2026
Framework
N/A
License
N/A
Stars
14
Summary
Vision-Language-Action (VLA) models have shown promise for generalizable robot control but struggle when adapting to new...
Abstract Summary
Key Points
- Proposes PriorVLA, a novel approach for vla in robotics.
- Addresses key limitations in existing methods through innovative architecture design.
- Demonstrates strong empirical results on standard benchmarks and real-world evaluations.
- Introduces novel training methodology that improves generalization and sample efficiency.
- Provides comprehensive analysis of failure modes and ablation studies.
Abstract
Vision-Language-Action (VLA) models have shown promise for generalizable robot control but struggle when adapting to new domains while preserving prior capabilities. This paper introduces PriorVLA, a method that preserves pre-trained priors during VLA fine-tuning through constrained optimization in the latent action space.
Links
- Paper (PDF): 2605.10925
- arXiv: 2605.10925
Related Papers
ALAM: Algebraically Consistent Latent Transitions for Vision-Language-Action Models
Zuojin Tang, Haoyun Liu, Xinyuan Chang et al. · arXiv · May 2026
Vision-Language-Action (VLA) models map visual observations and language instructions to robot actions through learned l...
Unified Noise Steering for Efficient Human-Guided VLA Adaptation
Junjie Lu, Xinyao Qin, Yuhua Jiang et al. · arXiv · May 2026
Adapting Vision-Language-Action (VLA) models to specific user preferences or task requirements typically requires expens...
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
Wenxuan Guo, Xiuwei Xu, Yichen Liu et al. · arXiv preprint · May 2026
AwareVLN introduces selective self-aware reasoning for VLN agents, triggering explicit spatial and progress analysis only at uncertain waypoints to improve robustness and explainability.
DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog