GesVLA: Gesture-Aware Vision-Language-Action Model
FeaturedWenxuan Guo, Ziyuan Li, Meng Zhang, Yichen Liu, Yimeng Dong, Chuxi Xu, Yunfei Wei, Ze Chen, Erjin Zhou, Jianjiang Feng · Tsinghua University · 2026
Framework
N/A
License
N/A
Stars
18
Summary
GesVLA augments standard VLA models with gesture awareness, enabling robots to interpret verbal instructions alongside human hand gestures for disambiguated manipulation.
Abstract Summary
Key Points
- First VLA model to explicitly integrate hand-gesture awareness.
- Gesture-Aware Embedding (GAE) fuses 2D hand poses with language tokens.
- Disambiguates object references in cluttered, multi-object scenes.
- Evaluated on Franka robot with real-world and simulated tasks.
- Supports gesture-only commands for noise-sensitive environments.
Additional Notes
Overview
- Paper (PDF): 2605.22812
- arXiv: 2605.22812
Related Papers
- OpenVLA: An Open-Source Vision-Language-Action Model
- RT-2: Vision-Language-Action Models
- RT-1: Robotics Transformer for Real-World Control at Scale
Related Papers
3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training
Jiaxin Shi, Xidong Zhang, Fucai Zhu et al. · arXiv preprint · Jun 2026
We propose a 3D-thinking-guided co-training framework that enables vision-language-action (VLA) models to perform 3D spatial reasoning implicitly during action prediction. Our core insight is that 3D geometry perception and 3D spatial reasoning are distinct capabilities that can be disentangled and ...
AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots
Jianli Sun, Bin Tian, Qiyao Zhang et al. · arXiv preprint · Jun 2026
Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-level Unmanned Aerial Vehicle (UAV) movement and end-effector-level arm manipulation differ substantially in action scale, dynamics, and control objectives. In this paper, we propose AIR-VL...
AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation
Mengfei Zhao, Dihong Huang, Yikai Tang et al. · arXiv preprint · Jul 2026
Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine a...
BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models
Zhongxi Chen, Yifan Han, Yanming Shao et al. · arXiv preprint · May 2026
Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulation. However, dexterous manipulation remains chall