ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, Davis Rempe · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for intera...
Abstract Summary
Key Points
- Generating realistic 3D human motions in real-time within interactive applications is key for ani...
- While recent offline motion generation approaches offer precise control via text and kinematic co...
- Conversely, existing online methods enable real-time synthesis but often sacrifice controllabilit...
- In this work, we introduce ARDY, a streaming generation framework that bridges this gap by enabli...
- ARDY employs a hybrid representation that combines explicit root features with a latent body embe...
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
|Authors: Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, Davis Rempe
|Venue: arXiv preprint | Year: 2026
|arXiv: 2607.08741v1
Abstract
Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows. In this work, we introduce ARDY, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints. ARDY employs a hybrid representation that combines explicit root features with a latent body embedding, balancing precise trajectory control with efficient generative learning. We propose a two-stage autoregressive transformer denoiser that features variable history context and supports conditioning on flexible, long-horizon kinematic constraints. By training on a large-scale motion capture dataset and being directly conditioned on text labels and kinematic constraints sampled from ground truth poses, ARDY natively learns controllable generation that supports online prompting and flexible long-horizon goals. Extensive evaluations on the HumanML3D benchmark and the large-scale, high-fidelity Bones Rigplay dataset demonstrate ARDY’s high motion quality and constraint adherence, validating the efficacy of our key architectural decisions. Finally, we demonstrate the method’s practical versatility through an interactive demo featuring dynamic text control, diverse keyframe pose constraints, path following, and interactive locomotion control via mouse and keyboard. Supplementary video results, code, and model releases can be found at https://research.nvidia.com/labs/sil/projects/ardy/.
Key Contributions
- Generating realistic 3D human motions in real-time within interactive applications is key for ani…
- While recent offline motion generation approaches offer precise control via text and kinematic co…
- Conversely, existing online methods enable real-time synthesis but often sacrifice controllabilit…
- In this work, we introduce ARDY, a streaming generation framework that bridges this gap by enabli…
- ARDY employs a hybrid representation that combines explicit root features with a latent body embe…
Topics
- locomotion
- reinforcement-learning
- control
- benchmark
Code & Data
No code repository linked in paper metadata.
BibTeX
@article{Zhao2026_260708741v1,
title = {ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation},
author = {Kaifeng Zhao and Mathis Petrovich and Haotian Zhang and Tingwu Wang and Siyu Tang and Davis Rempe},
year = {2026},
eprint = {2607.08741v1},
archivePrefix = {arXiv},
primaryClass = {cs.GR},
url = {https://arxiv.org/abs/2607.08741v1}
}
Related Papers
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation
Siyuan Ma, Boshi Zhang, Yutian Zhang et al. · arXiv preprint · Aug 2026
Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOW...
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Zhe Li, Zhenzhe Zhang, Yangyang Wei et al. · arXiv preprint · Aug 2026
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
Haoran Jiang, Jin Chen, Qingwen Bu et al. · arXiv preprint · Dec 2025
Humanoid robots require precise locomotion and dexterous manipulation to perform challenging loco-manipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This confines the robot to a limited workspace, preventing it from performing large...
A Browser-Native Digital Test Range for Benchmarking 4D Ocean-Glider Planning Algorithms
Edward Holmberg, Elias Ioup, Mahdi Abdelguerfi · arXiv preprint · Aug 2026
Repeated in-situ evaluation of ocean-glider planners requires scarce vehicles, operators, deployment and recovery resources, and ocean conditions that cannot be reset for competing algorithms. We present a guided, installation-free browser-native digital test range that transforms a selected regi...