Capability-Aware Arbitration for Semantic Intent-Based Shared Control
Zhaoda Du, Michael Bowman, Xiaoli Zhang · N/A · 2026
Framework
N/A
License
N/A
Stars
N/A
Summary
Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (...
Abstract Summary
Key Points
- Shared control often allocates robot authority based on confidence in inferred human intent, assu...
- When this assumption fails, high intent confidence can cause over-helping
- We present a capability-aware shared-control framework in which a vision-language model (VLM) inf...
- VLA capability confidence is estimated online from the dispersion and local instability of stocha...
- We design a nonlinear arbitration policy that combines Bayesian-filtered semantic-intent confiden...
Method
Figure 1 summarizes the proposed shared-control framework. The VLM infers the human’s semantic intent from the observed motion and scene, generates the corresponding instruction $z_{R}$, and provides semantic-intent confidence $\alpha$. The instruction conditions the VLA, which generates the autonomous command $\boldsymbol{q}_{R}$ and provides capability confidence $\beta$.
Experiments
Experiments used a six-joint LeRobot SO-101 arm [ 22 ] , two Nintendo Switch Joy-Cons, and top- and wrist-view RGB cameras (Fig. 4 ). Five instructions (3 pick-and-place and 2 stacking directions) were evaluated with cubes and LEGOs from the in distribution (ID) and out-of distribution (OOD). We collected 400 ID demonstrations (80 per instruction), using episode-disjoint training, validation, and test splits.
Moondream3 [ 23 ] was adapted using rank-8 LoRA with a frozen backbone to score five complete instructions.
Results
Figure 5 illustrates how $\beta$ adapts $w_{R}$ in ID and OOD pick-and-place examples (a,b) and an ID stacking example (c). Across all three examples, semantic intent-only arbitration can assign high $w_{R}$ when $\alpha$ is high despite low $\beta$. In (a), $\alpha$ remains relatively stable, while $\beta$ increases during the later transport-and-placement phase, allowing the proposed $w_{R}$ to rise accordingly.
Sources and Demonstrations
Figures are reproduced from Du et al. See the full paper for experimental details and the project page, when the paper links one, for demonstrations and videos.
Related Papers
A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies
Mathilde Kappel, Clémence Grislain, Mohamed Chetouani et al.
Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as ...
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al.
Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Zhe Li, Zhenzhe Zhang, Yangyang Wei et al.
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
NVIDIA, :, Johan Bjorck et al.
General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for...