发表机构
Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对执行器退化导致动作失真的问题,提出遥测感知动作修正(TeAR),利用机载遥测数据修正冻结策略的输出动作,在18个策略-任务对及物理机械臂上显著提升成功率。
AI 中文摘要
机器人操作策略通常在假设指令动作即使在运行数小时后仍能产生与训练时相同运动的条件下进行训练。真实硬件违反了这一假设,因为电机逐渐升温、接触时电流饱和、负载下电压下降,因此相同的策略动作可能产生更弱、延迟或更嘈杂的运动。这些条件已经由机载遥测数据(如关节温度、电机电流和电源电压)测量到,然而这些信号通常仅用于记录或安全检查,而非策略自适应。我们提出了遥测感知动作修正(TeAR),这是一种策略无关的方法,通过在其输出动作到达底层控制器之前进行修正,将冻结的操作策略转变为遥测条件策略。TeAR学习了一个轻量级Transformer,将提议的动作与实时执行器遥测数据结合,并放大、衰减或偏置各个动作分量。我们在涵盖8个策略族和5个操作任务的18个策略-任务对上评估了TeAR。在额外的配对评估中,当存在退化模型失配时,TeAR实现了31.8%的成功率,而基础策略为25.6%,假设模型逆为30.6%。在物理机械臂上,TeAR在加热条件下将成功率提高了10-15%,且无需在机器人上进行微调。
英文摘要
Robot manipulation policies are usually trained under the assumption that a commanded action produces the same motion as it did during training even after hours of operation. Real hardware violates this assumption as the motors gradually heat up, current saturates near contact, voltage sags under load, thus the same policy action can produce a weaker, delayed, or noisier motion. These conditions are already measured by onboard telemetry, such as joint temperature, motor current, and supply voltage, yet this signal is typically used only for logging or safety checks rather than policy adaptation. We introduce Telemetry-Aware Action Rectification (TeAR), a policy-agnostic method that turns a frozen manipulation policy into a telemetry-conditioned policy by rectifying its outgoing action before it reaches the low-level controller. TeAR learns a lightweight Transformer that combines the proposed action with live actuator telemetry and amplifies, damps, or biases individual action components. We evaluate TeAR across 18 policy-task pairs spanning 8 policy families and 5 manipulation tasks. In an additional paired evaluation with degradation-model mismatch, TeAR achieves 31.8% success, compared with 25.6% for the base policy and 30.6% for an assumed-model inverse. On a physical arm, TeAR improves success under heating by 10-15% without on-robot fine-tuning.
Comments13 pages, 15 figures