SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

He Zhu, Lusen Zhao, Kwan Man Cheng, Su Li, Katerina Fragkiadaki · N/A · 2026

Framework

N/A

License

N/A

Stars

N/A

Summary

Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines oft...

Abstract Summary

Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop controllers, scripted skill sequences, or task-specific programs. We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by exploring over Neural Interaction Skills (NIS): reusable, parameterized, closed-loop policies that expose learned physical interaction capabilities to a reasoning agent. Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and parameterizes NIS to interact with the environment, observes their outcomes, and generates verification, reflection, and memory to guide subsequent exploration. We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation and organize exploration as verifier-guided tree search, enabling the agent to discover successful long-horizon behaviors without relying on predetermined execution pipelines. SkillWeaver scales autonomously to 39.1K demonstrations across 14.1K scenes, which we distill into visuomotor policies. Across simulation benchmarks and real-world manipulation, training on SkillWeaver-generated experience substantially improves generalization to novel objects, spatial configurations, tasks, and environments, and enables zero- and few-shot sim-to-sim and sim-to-real transfer. Our results suggest agentic exploration over neural interaction skills as a scalable alternative for robot data generation.

Key Points

  • Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting r...
  • Simulation offers a scalable alternative, but existing data-generation pipelines often rely on op...
  • We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by ex...
  • Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and ...
  • We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation a...

SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

|Authors: He Zhu, Lusen Zhao, Kwan Man Cheng, Su Li, Katerina Fragkiadaki

|Venue: arXiv preprint | Year: 2026

|arXiv: 2609.36171

Abstract

Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop controllers, scripted skill sequences, or task-specific programs. We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by exploring over Neural Interaction Skills (NIS): reusable, parameterized, closed-loop policies that expose learned physical interaction capabilities to a reasoning agent. Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and parameterizes NIS to interact with the environment, observes their outcomes, and generates verification, reflection, and memory to guide subsequent exploration. We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation and organize exploration as verifier-guided tree search, enabling the agent to discover successful long-horizon behaviors without relying on predetermined execution pipelines. SkillWeaver scales autonomously to 39.1K demonstrations across 14.1K scenes, which we distill into visuomotor policies. Across simulation benchmarks and real-world manipulation, training on SkillWeaver-generated experience substantially improves generalization to novel objects, spatial configurations, tasks, and environments, and enables zero- and few-shot sim-to-sim and sim-to-real transfer. Our results suggest agentic exploration over neural interaction skills as a scalable alternative for robot data generation.

Key Contributions

  • Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting r…
  • Simulation offers a scalable alternative, but existing data-generation pipelines often rely on op…
  • We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by ex…
  • Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and …
  • We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation a…

Topics

  • manipulation
  • vla
  • reinforcement-learning
  • sim-to-real
  • control
  • learning-from-demonstration
  • benchmark

Code & Data

No code repository linked in paper metadata.

BibTeX

@article{Zhu2026_260936171,
  title     = {SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation},
  author    = {He Zhu and Lusen Zhao and Kwan Man Cheng and Su Li and Katerina Fragkiadaki},
  year      = {2026},
  eprint    = {2609.36171},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url       = {https://arxiv.org/abs/2609.36171}
}
Share

Related Papers

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies
arXiv preprint

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

Mathilde Kappel, Clémence Grislain, Mohamed Chetouani et al.

Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as ...

manipulation vision vla reinforcement-learning sim-to-real control learning-from-demonstration benchmark
PDF Advanced
Capability-Aware Arbitration for Semantic Intent-Based Shared Control
arXiv preprint

Capability-Aware Arbitration for Semantic Intent-Based Shared Control

Zhaoda Du, Michael Bowman, Xiaoli Zhang

Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
Code PDF Intermediate
GitHub ★ —
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
arXiv preprint

DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al.

Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
PDF Advanced
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
arXiv preprint

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Zhe Li, Zhenzhe Zhang, Yangyang Wei et al.

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...

manipulation locomotion vision vla reinforcement-learning control learning-from-demonstration benchmark
PDF Advanced