Native Video-Action Pretraining for Generalizable Robot Control
Qihang Zhang, Lin Li, Luyao Zhang, Shuai Yang, Yiming Luo, Shuaiting Li, Ruilin Wang, Junke Wang, Jiahao Shao, Gangwei Xu, Jiaming Zhou, Yishu Shen, Yudong Jin, Fangyi Xu, Shuailei Ma, Jiaqi Liao, Guanxing Lu, Zifan Shi, Yongkun Wen, Yujie Zhao, Weixuan Tang, Xinyang Wang, Chaojian Li, Jiapeng Zhu, Ka Leong Cheng, Nan Xue, Xing Zhu, Yujun Shen, Yinghao Xu · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments. To bridge this gap, we present LingBot-VA 2.0, a video-action foun...
Abstract Summary
Key Points
- The advent of video-action models offers a promising path for robot control
- Nevertheless, we argue that repurposing video generative models designed for digital content crea...
- To bridge this gap, we present LingBot-VA 2.0, a video-action foundation model built from the gro...
- Four core design principles showcase its evolution from LingBot-VA
- (1) Departing from traditional reconstruction-focused VAEs, we introduce a semantic visual-action...
Native Video-Action Pretraining for Generalizable Robot Control
|Authors: Qihang Zhang, Lin Li, Luyao Zhang, Shuai Yang, Yiming Luo, Shuaiting Li, Ruilin Wang, Junke Wang, Jiahao Shao, Gangwei Xu, Jiaming Zhou, Yishu Shen, Yudong Jin, Fangyi Xu, Shuailei Ma, Jiaqi Liao, Guanxing Lu, Zifan Shi, Yongkun Wen, Yujie Zhao, Weixuan Tang, Xinyang Wang, Chaojian Li, Jiapeng Zhu, Ka Leong Cheng, Nan Xue, Xing Zhu, Yujun Shen, Yinghao Xu
|Venue: arXiv preprint | Year: 2026
|arXiv: 2607.08639v1
Abstract
The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments. To bridge this gap, we present LingBot-VA 2.0, a video-action foundation model built from the ground up for embodiment. Four core design principles showcase its evolution from LingBot-VA. (1) Departing from traditional reconstruction-focused VAEs, we introduce a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning. (2) Given the strictly causal nature of temporal dynamics, we adopt a causal pretraining paradigm, training from scratch to circumvent the catastrophic forgetting that frequently occurs when adapting bidirectional architectures. (3) To meet the demands of high-frequency inference, our model employs a sparse MoE backbone, expanding model capacity without compromising efficiency. (4) Real-time closed-loop control is realized through an enhanced asynchronous inference scheme, which predicts future latents in parallel with action execution while re-grounding each rollout on the latest observation via learned forward dynamics. Real-world deployment validates LingBot-VA 2.0 as a robust foundation model, as evidenced by its few-shot generalization across complex manipulation tasks.
Key Contributions
- The advent of video-action models offers a promising path for robot control
- Nevertheless, we argue that repurposing video generative models designed for digital content crea…
- To bridge this gap, we present LingBot-VA 2.0, a video-action foundation model built from the gro…
- Four core design principles showcase its evolution from LingBot-VA
- (1) Departing from traditional reconstruction-focused VAEs, we introduce a semantic visual-action…
Topics
- manipulation
- vision
- reinforcement-learning
- control
Code & Data
No code repository linked in paper metadata.
BibTeX
@article{Zhang2026_260708639v1,
title = {Native Video-Action Pretraining for Generalizable Robot Control},
author = {Qihang Zhang and Lin Li and Luyao Zhang and Shuai Yang and Yiming Luo and Shuaiting Li and Ruilin Wang and Junke Wang and Jiahao Shao and Gangwei Xu and Jiaming Zhou and Yishu Shen and Yudong Jin and Fangyi Xu and Shuailei Ma and Jiaqi Liao and Guanxing Lu and Zifan Shi and Yongkun Wen and Yujie Zhao and Weixuan Tang and Xinyang Wang and Chaojian Li and Jiapeng Zhu and Ka Leong Cheng and Nan Xue and Xing Zhu and Yujun Shen and Yinghao Xu},
year = {2026},
eprint = {2607.08639v1},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2607.08639v1}
}
Related Papers
Decoding Task Progress from VLA Representations
Atiksh Bhardwaj, Edward Weiyi Duan, Prithwish Dan et al. · arXiv preprint · Aug 2026
Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools for understanding what these models represent internally or for monitoring them at runtime. Leveraging ideas from mechanistic interpretability, we...
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation
Siyuan Ma, Boshi Zhang, Yutian Zhang et al. · arXiv preprint · Aug 2026
Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOW...
Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators
Juan José García Cárdenas, Alperen Kenan, Hamidreza Raei et al. · arXiv preprint · Aug 2026
Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in challenging environments. In the nuclear industry, surface contact tasks such as swab sampling require precise path and force tracking, obstacle avoidance, and sustained operator attenti...
Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction
Thang Tran Viet, Thanh Nguyen Canh, Huy Uong Gia et al. · arXiv preprint · Sep 2026
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in real-world environments. This paper introduces a humanoid robot prototype designed as a flexible testbed for developing and integrating artificial intelligence (AI) modules in HRI tasks. The...