IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Featured

Shijie Lian, Bin Yu, Xiaopeng Lin, Zhaolong Shen, Laurence Tianruo Yang, Yurun Jin, Haishan Liu, Changti Wu, Hang Yuan, Cong Huang, Kai Chen · · 2026

Framework

N/A

License

N/A

Stars

8

Summary

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current o...

Abstract Summary

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines

Key Points

  • Introduces a new dataset or benchmark
  • Provides open-source code or data

Abstract

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines

Share

Related Papers

Scalable Behavior Cloning with Open Data, Training, and Evaluation

Scalable Behavior Cloning with Open Data, Training, and Evaluation

Arthur Allshire, Himanshu Gaurav Singh, Ritvik Singh et al. · arXiv · Jun 2026

We introduce ABC, a fully open-source stack for manipulation with behavior cloning. At its core is ABC-130K: the largest open-source teleoperation dataset to date, featuring 3,500 hours of data spanning over 130K episodes across 195 diverse tasks. Furthermore, we open-source o...

Manipulation VLA Models Reinforcement Learning Imitation Learning Sensing & Perception
Code PDF Intermediate
GitHub ★ 200 Code updated: Jun 2026
Hand-in-the-Loop: Improving Dexterous VLA via Seamless Interventional Correction

Hand-in-the-Loop: Improving Dexterous VLA via Seamless Interventional Correction

Zhuohang Li, Liqun Huang, Wei Xu et al. · arXiv · May 2026

Vision-Language-Action (VLA) models are prone to compounding errors in dexterous manipulation, where high-dimensional action spaces and contact-rich dynamics amplify small policy deviations over long horizons. While Interactive Imitation Learning (IIL) can refine policies through human takeover data...

Manipulation Imitation Learning VLA Models Sensing & Perception
PDF Advanced
No code repo Code updated: May 2026
LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Tao Lin, Yuxin Du, Yiran Mao et al. · arXiv · Jun 2026

Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visual-action supervision can dominate the comparatively sparse language-action signal. As a result, ...

Manipulation VLA Models Reinforcement Learning Imitation Learning Sensing & Perception
PDF Intermediate
No code repo Code updated: Jun 2026
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

Wen Ye, Peiyan Li, Tingyu Yuan et al. · arXiv · Jun 2026

Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major challenges remain unsolved: (1) reasoning can effectively improve the performance of the policy, but its scaling mechanism has seldom been studied; (2) historical i...

Manipulation VLA Models Reinforcement Learning Sensing & Perception
PDF Intermediate
No code repo Code updated: Jun 2026