Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation
Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA ...
Abstract Summary
Key Points
- Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in ...
- We investigate a neuro-symbolic framework that combines learned VLA control with explicit task gr...
- Task graphs encode action dependencies, valid transitions, and branch conditions, while memory ma...
- Together, these structures guide object selection, destination grounding, subgoal dispatch, and v...
- Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues
Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation
|Authors: Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger
|Venue: arXiv preprint | Year: 2026
|arXiv: 2609.05369v1
Abstract
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory maintains the active step, completed actions, textual context, and task-relevant visual evidence. Together, these structures guide object selection, destination grounding, subgoal dispatch, and verification of expected state transitions. Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues. To isolate their effect on policy learning, our initial study bypasses cross-view gaze transfer and directly annotates pseudo-gaze in robot-view teleoperation videos. The resulting guidance is used during VLA fine-tuning and inference. We study two long-horizon manipulation domains, workspace clearing and surgical-instrument handling, which require ordered execution, visually grounded decisions, and conditional branching. We evaluate correct-object and destination selection, subtask completion, task progress, step-order consistency, complete-task success, and procedural or execution mistakes. This work positions structured symbolic reasoning and demonstration-derived visual guidance as complementary mechanisms for reliable long-horizon VLA manipulation.
Key Contributions
- Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in …
- We investigate a neuro-symbolic framework that combines learned VLA control with explicit task gr…
- Task graphs encode action dependencies, valid transitions, and branch conditions, while memory ma…
- Together, these structures guide object selection, destination grounding, subgoal dispatch, and v…
- Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues
Topics
- manipulation
- vision
- vla
- reinforcement-learning
- control
- learning-from-demonstration
Code & Data
No code repository linked in paper metadata.
BibTeX
@article{Chavan2026_260905369v1,
title = {Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation},
author = {Vivek Chavan and Yahuan Shi and Oliver Heimann and Kevin Haninger and Jörg Krüger},
year = {2026},
eprint = {2609.05369v1},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.05369v1}
}
Related Papers
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al. · arXiv preprint · Jul 2026
Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Zhe Li, Zhenzhe Zhang, Yangyang Wei et al. · arXiv preprint · Aug 2026
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
NVIDIA, :, Johan Bjorck et al. · arXiv preprint · Mar 2025
General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for...
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
Haoran Jiang, Jin Chen, Qingwen Bu et al. · arXiv preprint · Dec 2025
Humanoid robots require precise locomotion and dexterous manipulation to perform challenging loco-manipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This confines the robot to a limited workspace, preventing it from performing large...