HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
Qiuxuan Feng, Jiale Yu, Jiaming Liu, Yueru Jia, Zhuangzhe Wu · · 2026
Framework
N/A
License
N/A
Stars
7
Summary
World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current...
Abstract Summary
Key Points
- Proposes HarmoWAM, a novel approach for foundation models in robotics.
- Addresses key limitations in existing methods through innovative architecture design.
- Demonstrates strong empirical results on standard benchmarks and real-world evaluations.
- Introduces novel training methodology that improves generalization and sample efficiency.
- Provides comprehensive analysis of failure modes and ablation studies.
Abstract
World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: generalizable but imprecise models learned from diverse data, or precise but specialized models fitted to specific tasks. This paper proposes HarmoWAM, an adaptive world action model that harmonizes both paradigms through dynamic task-aware modulation.
Links
- Paper (PDF): 2605.10942
- arXiv: 2605.10942
Related Papers
DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
Jusuk Lee, Seungjae Lee, Jonghun Shin et al. · arXiv preprint · May 2026
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
Christen Millerdurai, Shaoxiang Wang, Yaxu Xie et al. · SIGGRAPH 2026 · May 2026
Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where sensing must...
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
Xiaosong Jia, Bowen Yang, Zuhao Ge et al. · RSS 2026 · May 2026
Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLMs). Existing VLAs rely on end-to-end supervision to implicitly enable the action decoding process to learn...
Video
MOCHI: Motion Enhancement of Collaborative Human-object Interactions
Jiye Lee, Yonghun Choi, Jungdam Won · SIGGRAPH 2026 (Journal Track) · Jun 2026
Two-stage framework that enhances noisy multi-human object interaction (MHOI) data by generating plausible hand grasps and refining full-body motion via diffusion-based optimization.