Papers

Sorted by year (newest first)
PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

Sergey Arkhangelskiy · arXiv preprint · May 2026

Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with N <= 25 rollouts per condition, almost always without confidence intervals or

benchmarking vla manipulation evaluation
PDF Intermediate
No code repo Code updated: May 2026

Suggested Learning Path

Read these papers in order to build expertise in evaluation.

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6

…and 6 more papers.