Unified Noise Steering for Efficient Human-Guided VLA Adaptation

Unified Noise Steering for Efficient Human-Guided VLA Adaptation

Junjie Lu, Xinyao Qin, Yuhua Jiang, Kaixin Wang, Bolei Zhou · · 2026

Framework

N/A

License

N/A

Stars

0

Summary

Adapting Vision-Language-Action (VLA) models to specific user preferences or task requirements typically requires expens...

Abstract Summary

Adapting Vision-Language-Action (VLA) models to specific user preferences or task requirements typically requires expensive fine-tuning. This paper introduces Unified Noise Steering, a lightweight method that guides VLA model behavior by steering the diffusion noise process during inference rather t... The method demonstrates significant improvements over existing approaches, providing both theoretical insights and practical benefits for real-world deployment. Comprehensive experiments validate the effectiveness of the proposed approach across diverse scenarios and task settings.

Key Points

  • Proposes Unified Noise Steering for Efficient Human-Guided VLA Adaptation, a novel approach for vla in robotics.
  • Addresses key limitations in existing methods through innovative architecture design.
  • Demonstrates strong empirical results on standard benchmarks and real-world evaluations.
  • Introduces novel training methodology that improves generalization and sample efficiency.
  • Provides comprehensive analysis of failure modes and ablation studies.

Abstract

Adapting Vision-Language-Action (VLA) models to specific user preferences or task requirements typically requires expensive fine-tuning. This paper introduces Unified Noise Steering, a lightweight method that guides VLA model behavior by steering the diffusion noise process during inference rather than modifying model weights.

Share

Related Papers

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog

foundation-models perception manipulation representation-learning vla
PDF Intermediate
No code repo Code updated: May 2026