Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
FeaturedIsmail Geles, Leonard Bauersfeld, Markus Wulfmeier, Davide Scaramuzza · University of Zurich, Google DeepMind · 2026
Summary
The first demonstration of superhuman, safe, multi-agent drone racing using a MARL framework trained in simulation and transferred to real Crazyflie nano-drones.
Abstract Summary
Key Points
- First demonstration of superhuman safe multi-agent drone racing.
- Multi-agent RL trained in simulation transfers to real Crazyflie drones.
- Safety-conditioned action space balances aggression and collision avoidance.
- Outperforms human world champion lap times with zero collisions.
- Uses photorealistic sim-to-real transfer for outdoor deployment.
Additional Notes
Overview
- Paper (PDF): 2605.22748
- arXiv: 2605.22748
Related Papers
- Agile but Safe: Learning Collision-Free High-Speed Quadrupedal Maneuvers
- Eureka: Neural Network Synthesis via LLMs
- Fast-LIO 2: Fast Direct LiDAR-Inertial Odometry
Related Papers
Scout-Assisted Planning for Heterogeneous Robot Teams under Partially Known Environments
Hoang-Dung Bui, Abhish Khanal, Raihan Islam Arnob et al. · arXiv preprint · May 2026
SAP pairs aerial scouts with ground robots to gather environmental information ahead of time, reducing mission time by 40% through information-theoretic POMDP planning.
Learning Agile Flight in the Wild
Elia Kaufmann, Mathias Gehrig, Philipp Foehn et al. · Science Robotics · Feb 2023
An end-to-end neural drone controller trained entirely in simulation flies acrobatic maneuvers in the real world with zero real-world fine-tuning, enabled by domain randomization and reinforcement learning.
AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots
Jianli Sun, Bin Tian, Qiyao Zhang et al. · arXiv preprint · Jun 2026
Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-level Unmanned Aerial Vehicle (UAV) movement and end-effector-level arm manipulation differ substantially in action scale, dynamics, and control objectives. In this paper, we propose AIR-VL...
CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning
Hexian Ni, Tao Lu, Yinghao Cai · ICML 2026 · Jul 2026
Reward design remains a central challenge in reinforcement learning (RL). Hand-crafted rewards are often difficult to specify and may lead to suboptimal policies, while learned rew...