Capability-Aware Arbitration for Semantic Intent-Based Shared Control

Capability-Aware Arbitration for Semantic Intent-Based Shared Control

Zhaoda Du, Michael Bowman, Xiaoli Zhang · N/A · 2026

Framework

N/A

License

N/A

Stars

N/A

Summary

Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (...

Abstract Summary

Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (VLM) infers human intent and provides semantic-intent confidence, while a vision-language-action (VLA) policy generates autonomous actions. VLA capability confidence is estimated online from the dispersion and local instability of stochastic action trajectories. We design a nonlinear arbitration policy that combines Bayesian-filtered semantic-intent confidence with VLA capability confidence through a sigmoid mapping to adapt robot authority. Our evaluation combined VLM/VLA confidence assessment with a study involving 12 participants performing pick-and-place and bidirectional stacking under in-distribution and out-of-distribution conditions. The proposed method achieved the highest task success rate (92%), compared with manual teleoperation (83%), intent-only arbitration (44%), and fixed equal-weight blending (10%). It also achieved higher control friendliness and lower authority-weighted disagreement than both shared-control baselines. These results demonstrate the benefit of incorporating VLA capability into authority allocation to mitigate over-helping and improve shared-control performance.

Key Points

  • Shared control often allocates robot authority based on confidence in inferred human intent, assu...
  • When this assumption fails, high intent confidence can cause over-helping
  • We present a capability-aware shared-control framework in which a vision-language model (VLM) inf...
  • VLA capability confidence is estimated online from the dispersion and local instability of stocha...
  • We design a nonlinear arbitration policy that combines Bayesian-filtered semantic-intent confiden...

Method

Figure 1 summarizes the proposed shared-control framework. The VLM infers the human’s semantic intent from the observed motion and scene, generates the corresponding instruction $z_{R}$, and provides semantic-intent confidence $\alpha$. The instruction conditions the VLA, which generates the autonomous command $\boldsymbol{q}_{R}$ and provides capability confidence $\beta$.

Method — Capability-Aware Arbitration for Semantic Intent-Based Shared Control
Fig. 2: Bayesian intent filtering and capability-confidence mapping. (a) Unfiltered predictions (top) switch 24 times, whereas filtered predictions (bottom) switch only 2 times. The Bayesian transition probabilities are $P(z\mid z)=0.90$ and $P(z\mid z^{\prime})=0.025$ for $z\neq z^{\prime}$. (b) Mapping from trajectory dispersion $D_{t}$ and local instability $A_{t}$ to capability confidence $\beta_{t}$. Colors and solid contours indicate $\beta_{t}$. The dashed line marks $D_{t}/D_{0}+A_{t}/A_{0}=2$, beyond which $\beta_{t}=0$. $D_{0}=12.17\,\mathrm{mm}$ and $A_{0}=6.91\,\mathrm{mm}$ which are the 95th percentiles computed from 200 reference observations.

Experiments

Experiments used a six-joint LeRobot SO-101 arm [ 22 ] , two Nintendo Switch Joy-Cons, and top- and wrist-view RGB cameras (Fig. 4 ). Five instructions (3 pick-and-place and 2 stacking directions) were evaluated with cubes and LEGOs from the in distribution (ID) and out-of distribution (OOD). We collected 400 ID demonstrations (80 per instruction), using episode-disjoint training, validation, and test splits.

Moondream3 [ 23 ] was adapted using rank-8 LoRA with a frozen backbone to score five complete instructions.

Experiments — Capability-Aware Arbitration for Semantic Intent-Based Shared Control
Fig. 3: Effect of sigmoid mapping on robot control authority. (a) Direct multiplication exhibits multiplicative attenuation: $\alpha\beta$ is smaller than either confidence when $0<\alpha,\beta<1$; for example, $\alpha=\beta=0.7$ yields $w_{R}=0.49$. (b) The proposed mapping retains the multiplicative confidence gate but applies $w_{R}=\sigma\!\left(\kappa(\alpha\beta-\tau)\right)$, ($\tau=0.40$, $\kappa=5.0$), to sharpen the authority transition near the threshold and mitigate attenuation. (c) Difference between the two mappings, $\Delta w_{R}=w_{R}^{\mathrm{sigmoid}}-\alpha\beta$.

Results

Figure 5 illustrates how $\beta$ adapts $w_{R}$ in ID and OOD pick-and-place examples (a,b) and an ID stacking example (c). Across all three examples, semantic intent-only arbitration can assign high $w_{R}$ when $\alpha$ is high despite low $\beta$. In (a), $\alpha$ remains relatively stable, while $\beta$ increases during the later transport-and-placement phase, allowing the proposed $w_{R}$ to rise accordingly.

Results — Capability-Aware Arbitration for Semantic Intent-Based Shared Control
Fig. 4: Task design for the human study and experimental setup

Sources and Demonstrations

Figures are reproduced from Du et al. See the full paper for experimental details and the project page, when the paper links one, for demonstrations and videos.

Share

Related Papers

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies
arXiv preprint

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

Mathilde Kappel, Clémence Grislain, Mohamed Chetouani et al.

Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as ...

manipulation vision vla reinforcement-learning sim-to-real control learning-from-demonstration benchmark
PDF Advanced
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
arXiv preprint

DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

Yunchao Yao, Zhuxiu Xu, Tianqi Zhang et al.

Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embod...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
PDF Advanced
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
arXiv preprint

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Zhe Li, Zhenzhe Zhang, Yangyang Wei et al.

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action mode...

manipulation locomotion vision vla reinforcement-learning control learning-from-demonstration benchmark
PDF Advanced
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
arXiv preprint

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

NVIDIA, :, Johan Bjorck et al.

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
Code PDF Advanced
GitHub ★ —