HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Qiuxuan Feng, Jiale Yu, Jiaming Liu, Yueru Jia, Zhuangzhe Wu · · 2026

Framework

N/A

License

N/A

Stars

7

Summary

World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current...

Abstract Summary

World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: generalizable but imprecise models learned from diverse data, or precise but specialized models fitted to specific tasks. This paper proposes ... The method demonstrates significant improvements over existing approaches, providing both theoretical insights and practical benefits for real-world deployment. Comprehensive experiments validate the effectiveness of the proposed approach across diverse scenarios and task settings.

Key Points

  • Proposes HarmoWAM, a novel approach for foundation models in robotics.
  • Addresses key limitations in existing methods through innovative architecture design.
  • Demonstrates strong empirical results on standard benchmarks and real-world evaluations.
  • Introduces novel training methodology that improves generalization and sample efficiency.
  • Provides comprehensive analysis of failure modes and ablation studies.

Abstract

World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: generalizable but imprecise models learned from diverse data, or precise but specialized models fitted to specific tasks. This paper proposes HarmoWAM, an adaptive world action model that harmonizes both paradigms through dynamic task-aware modulation.

Share

Related Papers

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog

foundation-models perception manipulation representation-learning vla
PDF Intermediate
No code repo Code updated: May 2026
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

Christen Millerdurai, Shaoxiang Wang, Yaxu Xie et al. · SIGGRAPH 2026 · May 2026

Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where sensing must...

manipulation manipulation simulation
Code PDF Intermediate
Code ★ 0 May 2026