Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts
Zhen Sun, Yongjian Guo, Haoran Sun, Luqiao Wang, Wei Lu, Jiachi Ji, Shengzhe Ji, Junwu Xiong, Zhijun Meng · University of Warwick, University of Cambridge, HIT · 2026
Summary
A lightweight runtime verification layer that preemptively prunes unsafe VLA and world-model actions before they reach robot hardware, eliminating collision and droppage failures.
Abstract Summary
Key Points
- Runtime verification layer for VLA and world-model safety.
- Formal spatio-temporal task specifications for embodied tasks.
- Learned monitor preemptively prunes unsafe action sequences.
- Minimal overhead under 5 ms per action, deployable on embedded compute.
- Eliminates collision and droppage failures on household mobile manipulator.
Additional Notes
Overview
- Paper (PDF): 2605.22446
- arXiv: 2605.22446
Related Papers
- SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation
- Guide, Think and Act: Interactive Embodied Reasoning in VLA
- RT-2: Vision-Language-Action Models
Related Papers
Uncertainty Quantification for Flow-Based Vision-Language-Action Models
Ralf Römer, Maximilian Seeliger, Saida Liu et al. · RSS 2026 — Best Paper Award · Jun 2026
Quantifies epistemic uncertainty in flow-matching VLAs using velocity-field disagreement (VFD) across a small ensemble, enabling failure detection at deployment and sample-efficient active fine-tuning (SAVE).
ALAM: Algebraically Consistent Latent Transitions for Vision-Language-Action Models
Zuojin Tang, Haoyun Liu, Xinyuan Chang et al. · arXiv · May 2026
Vision-Language-Action (VLA) models map visual observations and language instructions to robot actions through learned l...
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
Wenxuan Guo, Xiuwei Xu, Yichen Liu et al. · arXiv preprint · May 2026
AwareVLN introduces selective self-aware reasoning for VLN agents, triggering explicit spatial and progress analysis only at uncertain waypoints to improve robustness and explainability.
DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog