PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models

Xinyu Guo, Bin Xie, Wei Chai, Xianchi Deng, Ruoxi Jia · · 2026

Framework

N/A

License

N/A

Stars

14

Summary

Vision-Language-Action (VLA) models have shown promise for generalizable robot control but struggle when adapting to new...

Abstract Summary

Vision-Language-Action (VLA) models have shown promise for generalizable robot control but struggle when adapting to new domains while preserving prior capabilities. This paper introduces PriorVLA, a method that preserves pre-trained priors during VLA fine-tuning through constrained optimization in ... The method demonstrates significant improvements over existing approaches, providing both theoretical insights and practical benefits for real-world deployment. Comprehensive experiments validate the effectiveness of the proposed approach across diverse scenarios and task settings.

Key Points

  • Proposes PriorVLA, a novel approach for vla in robotics.
  • Addresses key limitations in existing methods through innovative architecture design.
  • Demonstrates strong empirical results on standard benchmarks and real-world evaluations.
  • Introduces novel training methodology that improves generalization and sample efficiency.
  • Provides comprehensive analysis of failure modes and ablation studies.

Abstract

Vision-Language-Action (VLA) models have shown promise for generalizable robot control but struggle when adapting to new domains while preserving prior capabilities. This paper introduces PriorVLA, a method that preserves pre-trained priors during VLA fine-tuning through constrained optimization in the latent action space.

Share

Related Papers

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog

foundation-models perception manipulation representation-learning vla
PDF Intermediate
No code repo Code updated: May 2026