FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference
Zekai Li, Jiaming Tang, Zhijian Liu · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable asynchronous execution. This challenge is particularly pronounced in flow-matching-based VLA models, where action de...
Abstract Summary
Key Points
- Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet thei...
- This challenge is particularly pronounced in flow-matching-based VLA models, where action decodin...
- While efficient inference methods improve control frequency and asynchronous methods reduce execu...
- We introduce \textbf{FlashVLA}, a streaming action decoding framework that addresses both challen...
- FlashVLA maintains a streaming action buffer with multiple chunks at different noise levels and d...
FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference
|Authors: Zekai Li, Jiaming Tang, Zhijian Liu
|Venue: arXiv preprint | Year: 2026
|arXiv: 2608.27384v1
Abstract
Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable asynchronous execution. This challenge is particularly pronounced in flow-matching-based VLA models, where action decoding requires multiple iterative steps conditioned on the VLM context. While efficient inference methods improve control frequency and asynchronous methods reduce execution idle time, existing approaches often fail to jointly achieve low-latency inference and accurate, temporally consistent asynchronous execution. We introduce \textbf{FlashVLA}, a streaming action decoding framework that addresses both challenges in a unified formulation. FlashVLA maintains a streaming action buffer with multiple chunks at different noise levels and decodes them using chunk-wise causal attention. This design allows FlashVLA to produce one executable action chunk per inference step. Moreover, its chunk-wise autoregressive formulation implicitly preserves action continuity, enabling smooth asynchronous execution without extra future-state conditioning. Across extensive simulated and real-world experiments, FlashVLA substantially improves inference speed while maintaining strong task performance. It can achieve $\geq$30,Hz control frequency on a single GPU with smooth asynchronous inference in real-world deployment.
Key Contributions
- Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet thei…
- This challenge is particularly pronounced in flow-matching-based VLA models, where action decodin…
- While efficient inference methods improve control frequency and asynchronous methods reduce execu…
- We introduce \textbf{FlashVLA}, a streaming action decoding framework that addresses both challen…
- FlashVLA maintains a streaming action buffer with multiple chunks at different noise levels and d…
Topics
- manipulation
- vision
- vla
- reinforcement-learning
- control
Code & Data
- GitHub repository: https://github.com/z-lab/flashvla.git
BibTeX
@article{Li2026_260827384v1,
title = {FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference},
author = {Zekai Li and Jiaming Tang and Zhijian Liu},
year = {2026},
eprint = {2608.27384v1},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2608.27384v1}
}
Related Papers
Decoding Task Progress from VLA Representations
Atiksh Bhardwaj, Edward Weiyi Duan, Prithwish Dan et al. · arXiv preprint · Aug 2026
Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools for understanding what these models represent internally or for monitoring them at runtime. Leveraging ideas from mechanistic interpretability, we...
Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction
Thang Tran Viet, Thanh Nguyen Canh, Huy Uong Gia et al. · arXiv preprint · Sep 2026
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in real-world environments. This paper introduces a humanoid robot prototype designed as a flexible testbed for developing and integrating artificial intelligence (AI) modules in HRI tasks. The...
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al. · arXiv preprint · Jul 2026
Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
Junfeng Li, Junjie He, Zhide Zhong et al. · arXiv preprint · Aug 2026
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse v...