Test-Time Adaptation of Manipulation Policies Under Actuator Degradation

Test-Time Adaptation of Manipulation Policies Under Actuator Degradation

Som Sagar, Ransalu Senanayake · N/A · 2026

Framework

N/A

License

N/A

Stars

N/A

Summary

Robot manipulation policies are usually trained under the assumption that a commanded action produces the same motion as it did during training even after hours of operation. Real hardware violates this assumption as the motors gradually heat up, current saturates near contact, voltage sags under...

Abstract Summary

Robot manipulation policies are usually trained under the assumption that a commanded action produces the same motion as it did during training even after hours of operation. Real hardware violates this assumption as the motors gradually heat up, current saturates near contact, voltage sags under load, thus the same policy action can produce a weaker, delayed, or noisier motion. These conditions are already measured by onboard telemetry, such as joint temperature, motor current, and supply voltage, yet this signal is typically used only for logging or safety checks rather than policy adaptation. We introduce Telemetry-Aware Action Rectification (TeAR), a policy-agnostic method that turns a frozen manipulation policy into a telemetry-conditioned policy by rectifying its outgoing action before it reaches the low-level controller. TeAR learns a lightweight Transformer that combines the proposed action with live actuator telemetry and amplifies, damps, or biases individual action components. We evaluate TeAR across 18 policy-task pairs spanning 8 policy families and 5 manipulation tasks. In an additional paired evaluation with degradation-model mismatch, TeAR achieves 31.8% success, compared with 25.6% for the base policy and 30.6% for an assumed-model inverse. On a physical arm, TeAR improves success under heating by 10-15% without on-robot fine-tuning.

Key Points

  • Robot manipulation policies are usually trained under the assumption that a commanded action prod...
  • Real hardware violates this assumption as the motors gradually heat up, current saturates near co...
  • These conditions are already measured by onboard telemetry, such as joint temperature, motor curr...
  • We introduce Telemetry-Aware Action Rectification (TeAR), a policy-agnostic method that turns a f...
  • TeAR learns a lightweight Transformer that combines the proposed action with live actuator teleme...

Multi-Channel Degradation Model

We formalize deployment-time actuator degradation as a telemetry-conditioned perturbation of the policy action, giving controlled and stress conditions in simulation. Given per-joint temperature, current, and voltage measurements, our degradation wrapper computes a capacity factor $\rho_{c}(x_{c,j})\in(0,1]$ for each channel $c\in\{T,C,V\}$ and joint $j$.

A factor of one represents nominal capacity; smaller values represent reduced capacity under stress. The telemetry-dependent curves in Fig.

Multi-Channel Degradation Model — Test-Time Adaptation of Manipulation Policies Under Actuator Degradation
Fig. 5 : Environment setup for real (left-2) and sim (right-2).

TeAR Framework

Telemetry-Aware Action Rectification (TeAR) wraps a manipulation policy with a lightweight adapter that reads the three telemetry channels of Section III and corrects its outgoing action. The adapter is inserted between the base policy and the low-level controller, producing an action of the same dimensionality as the base command without changing the policy internals. Setup.

A frozen base policy maps task observations $o$ to a normalized action $a_{b}=\pi_{b}(o)\in[-1,1]^{d_{a}}$.

TeAR Framework — Test-Time Adaptation of Manipulation Policies Under Actuator Degradation
Fig. 6 : Stressed-condition $\Delta$ over the frozen base for five recipes with the TeAR architecture fixed.

Experiments

We organize the evaluation around five research questions. RQ1. Effectiveness: Does TeAR improve stressed-condition success across policy classes while preserving nominal behavior? RQ2. Comparison against baselines: How does it compare with robustness methods, alternative adapters, and training recipes? RQ3. Architectural attribution: Which components matter, and how does the trained adapter modify actions? RQ4.

Experiments — Test-Time Adaptation of Manipulation Policies Under Actuator Degradation
Fig. 7 : Stressed success gain sweeping the additive-residual cap $\alpha$ and the multiplicative gain $\gamma_{\max}$. The selected $(\alpha,\gamma_{\max})=(0.30,0.75)$ is outlined.

Sources and Demonstrations

Figures are reproduced from Sagar et al. See the full paper for experimental details and the project page, when the paper links one, for demonstrations and videos.

Share

Related Papers

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies
arXiv preprint

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

Mathilde Kappel, Clémence Grislain, Mohamed Chetouani et al.

Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as ...

manipulation vision vla reinforcement-learning sim-to-real control learning-from-demonstration benchmark
PDF Advanced
Capability-Aware Arbitration for Semantic Intent-Based Shared Control
arXiv preprint

Capability-Aware Arbitration for Semantic Intent-Based Shared Control

Zhaoda Du, Michael Bowman, Xiaoli Zhang

Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (...

manipulation vision vla reinforcement-learning control learning-from-demonstration benchmark
Code PDF Intermediate
GitHub ★ —
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation
arXiv preprint

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

Siyuan Ma, Boshi Zhang, Yutian Zhang et al.

Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOW...

manipulation locomotion vision reinforcement-learning control benchmark
PDF Intermediate
Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators
arXiv preprint

Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators

Juan José García Cárdenas, Alperen Kenan, Hamidreza Raei et al.

Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in challenging environments. In the nuclear industry, surface contact tasks such as swab sampling require precise path and force tracking, obstacle avoidance, and sustained operator attenti...

manipulation vision reinforcement-learning control learning-from-demonstration tactile benchmark
PDF Intermediate