RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

Huashuo Lei, Wenxuan Song, Huarui Zhang, Jieyuan Pei, Ziyue Zheng · · 2026

Framework

N/A

License

N/A

Stars

76

Summary

Long-horizon robotic tasks require agents to remember and reason about past observations and actions. This paper present...

Abstract Summary

Long-horizon robotic tasks require agents to remember and reason about past observations and actions. This paper presents RoboMemArena, a comprehensive benchmark suite designed to evaluate robotic memory across multiple dimensions including spatial memory, temporal reasoning, object permanence, and ... The method demonstrates significant improvements over existing approaches, providing both theoretical insights and practical benefits for real-world deployment. Comprehensive experiments validate the effectiveness of the proposed approach across diverse scenarios and task settings.

Key Points

  • Proposes RoboMemArena, a novel approach for perception in robotics.
  • Addresses key limitations in existing methods through innovative architecture design.
  • Demonstrates strong empirical results on standard benchmarks and real-world evaluations.
  • Introduces novel training methodology that improves generalization and sample efficiency.
  • Provides comprehensive analysis of failure modes and ablation studies.

Abstract

Long-horizon robotic tasks require agents to remember and reason about past observations and actions. This paper presents RoboMemArena, a comprehensive benchmark suite designed to evaluate robotic memory across multiple dimensions including spatial memory, temporal reasoning, object permanence, and task-dependent recall.

Share

Related Papers

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog

foundation-models perception manipulation representation-learning vla
PDF Intermediate
No code repo Code updated: May 2026
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

Christen Millerdurai, Shaoxiang Wang, Yaxu Xie et al. · SIGGRAPH 2026 · May 2026

Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where sensing must...

manipulation manipulation simulation
Code PDF Intermediate
Code ★ 0 May 2026
SI-Diff: A Framework for Learning Search and High-Precision Insertion with a Force-Domain Diffusion Policy

SI-Diff: A Framework for Learning Search and High-Precision Insertion with a Force-Domain Diffusion Policy

Yibo Liu, Stanko Oparnica, Simon Shewchun-Jakaitis et al. · arXiv · May 2026

Contact-rich manipulation is fundamental in robotics but poses significant challenges due to uncertainties in relative poses, such as misalignments and small clearances in peg-in-hole tasks. Existing approaches typically address search and...

manipulation tactile manipulation
Code PDF Intermediate
Code ★ 0 May 2026
3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training

Jiaxin Shi, Xidong Zhang, Fucai Zhu et al. · arXiv preprint · Jun 2026

We propose a 3D-thinking-guided co-training framework that enables vision-language-action (VLA) models to perform 3D spatial reasoning implicitly during action prediction. Our core insight is that 3D geometry perception and 3D spatial reasoning are distinct capabilities that can be disentangled and ...

sim-to-real reinforcement-learning vision vla manipulation
PDF Advanced
No code repo Jun 2026