GesVLA: Gesture-Aware Vision-Language-Action Model
Wenxuan Guo, Ziyuan Li, Meng Zhang et al. · arXiv preprint · May 2026
GesVLA augments standard VLA models with gesture awareness, enabling robots to interpret verbal instructions alongside human hand gestures for disambiguated manipulation.
vla manipulation hand-tracking
GitHub ★ 18 Code updated: May 2026