LANTERN: A Closed-Loop Benchmark for VLM-Based Cooperative Driving with Temporally Grounded Warnings
Yongshuo Liu, Xu Gao, Morui Zhu, Yongqi Zhu, Qi Chen, Deyuan Qu, Song Fu, Qing Yang · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
We present LANTERN, a closed-loop benchmark for temporally grounded cooperative warnings. LANTERN separates warning onset, hazard onset, warning termination, and post-hazard recovery, and evaluates each physical event under matched warning and no-warning executions so that the warning's contribut...
Abstract Summary
Key Points
- We present LANTERN, a closed-loop benchmark for temporally grounded cooperative warnings
- LANTERN separates warning onset, hazard onset, warning termination, and post-hazard recovery, and...
- The benchmark spans six safety-critical scenario families and provides 3,272 sequences with 236,3...
- Each hazard route is evaluated under the warning and no-warning conditions, while its no-hazard c...
- We further introduce the Cooperative Unified Score (CUS), a safety-gated metric that jointly rewa...
LANTERN: A Closed-Loop Benchmark for VLM-Based Cooperative Driving with Temporally Grounded Warnings
|Authors: Yongshuo Liu, Xu Gao, Morui Zhu, Yongqi Zhu, Qi Chen, Deyuan Qu, Song Fu, Qing Yang
|Venue: arXiv preprint | Year: 2026
|arXiv: 2609.06368v1
Abstract
We present LANTERN, a closed-loop benchmark for temporally grounded cooperative warnings. LANTERN separates warning onset, hazard onset, warning termination, and post-hazard recovery, and evaluates each physical event under matched warning and no-warning executions so that the warning’s contribution is measured in isolation rather than confounded with onboard vision. The benchmark spans six safety-critical scenario families and provides 3,272 sequences with 236,309 frames for training, together with 120 matched route pairs for closed-loop evaluation. Each hazard route is evaluated under the warning and no-warning conditions, while its no-hazard control penalizes unconditional braking. We further introduce the Cooperative Unified Score (CUS), a safety-gated metric that jointly rewards route progress, anticipation, clearance, and recovery. Fine-tuning a representative VLM driving model raises CUS from 34.6 without warnings to 75.5 with them, demonstrating both the value of cooperative warnings and the discriminative power of the paired protocol. All resources will be made publicly available.
Key Contributions
- We present LANTERN, a closed-loop benchmark for temporally grounded cooperative warnings
- LANTERN separates warning onset, hazard onset, warning termination, and post-hazard recovery, and…
- The benchmark spans six safety-critical scenario families and provides 3,272 sequences with 236,3…
- Each hazard route is evaluated under the warning and no-warning conditions, while its no-hazard c…
- We further introduce the Cooperative Unified Score (CUS), a safety-gated metric that jointly rewa…
Topics
- vision
- vla
- control
- benchmark
Code & Data
No code repository linked in paper metadata.
BibTeX
@article{Liu2026_260906368v1,
title = {LANTERN: A Closed-Loop Benchmark for VLM-Based Cooperative Driving with Temporally Grounded Warnings},
author = {Yongshuo Liu and Xu Gao and Morui Zhu and Yongqi Zhu and Qi Chen and Deyuan Qu and Song Fu and Qing Yang},
year = {2026},
eprint = {2609.06368v1},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.06368v1}
}
Related Papers
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al. · arXiv preprint · Jul 2026
Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
Junfeng Li, Junjie He, Zhide Zhong et al. · arXiv preprint · Aug 2026
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse v...
FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects
Chenhuan Liu, Yi Xu, Feng Wu et al. · arXiv preprint · Sep 2026
Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-object manipulation, where policies must track changing s...
Planning-Oriented End-to-End Autonomous Driving: Architectures, Evaluation, and Emerging Paradigms
Yanchen Guan, Xingcheng Liu, Bin Rao et al. · arXiv preprint · Aug 2026
End-to-end autonomous driving has evolved from camera-to-control regression toward planning-oriented systems that use structured representations, trajectory-level outputs, and increasingly realistic evaluation protocols. This survey reviews this transition across behavior cloning, conditional imi...