SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation
He Zhu, Lusen Zhao, Kwan Man Cheng, Su Li, Katerina Fragkiadaki · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines oft...
Abstract Summary
Key Points
- Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting r...
- Simulation offers a scalable alternative, but existing data-generation pipelines often rely on op...
- We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by ex...
- Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and ...
- We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation a...
SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation
|Authors: He Zhu, Lusen Zhao, Kwan Man Cheng, Su Li, Katerina Fragkiadaki
|Venue: arXiv preprint | Year: 2026
|arXiv: 2609.36171
Abstract
Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop controllers, scripted skill sequences, or task-specific programs. We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by exploring over Neural Interaction Skills (NIS): reusable, parameterized, closed-loop policies that expose learned physical interaction capabilities to a reasoning agent. Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and parameterizes NIS to interact with the environment, observes their outcomes, and generates verification, reflection, and memory to guide subsequent exploration. We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation and organize exploration as verifier-guided tree search, enabling the agent to discover successful long-horizon behaviors without relying on predetermined execution pipelines. SkillWeaver scales autonomously to 39.1K demonstrations across 14.1K scenes, which we distill into visuomotor policies. Across simulation benchmarks and real-world manipulation, training on SkillWeaver-generated experience substantially improves generalization to novel objects, spatial configurations, tasks, and environments, and enables zero- and few-shot sim-to-sim and sim-to-real transfer. Our results suggest agentic exploration over neural interaction skills as a scalable alternative for robot data generation.
Key Contributions
- Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting r…
- Simulation offers a scalable alternative, but existing data-generation pipelines often rely on op…
- We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by ex…
- Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and …
- We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation a…
Topics
- manipulation
- vla
- reinforcement-learning
- sim-to-real
- control
- learning-from-demonstration
- benchmark
Code & Data
No code repository linked in paper metadata.
BibTeX
@article{Zhu2026_260936171,
title = {SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation},
author = {He Zhu and Lusen Zhao and Kwan Man Cheng and Su Li and Katerina Fragkiadaki},
year = {2026},
eprint = {2609.36171},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.36171}
}
Related Papers
A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies
Mathilde Kappel, Clémence Grislain, Mohamed Chetouani et al.
Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as ...
Capability-Aware Arbitration for Semantic Intent-Based Shared Control
Zhaoda Du, Michael Bowman, Xiaoli Zhang
Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (...
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al.
Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Zhe Li, Zhenzhe Zhang, Yangyang Wei et al.
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...