发表机构
National Institute of AIST; Embodied AI Research Team(日本国立AIST研究所; 具身人工智能研究团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出ORPA框架,通过添加轻量反馈条件化模块实现机器人操控策略的在线残差适配,无需重训练即可纠正执行误差,在ALOHA平台的精度敏感操控任务上提升了成功率与扰动恢复能力。
AI 中文摘要
通过模仿学习训练的机器人操控策略,例如带Transformer的动作分块(ACT),在理想条件下可实现优异性能,但通常对微小执行误差和分布偏移较为敏感。纠正此类故障通常需要数据集聚合和全策略重训练,这在计算上成本高昂且不适合实时部署。本研究提出在线残差策略适配(ORPA)框架,该框架无需修改底层策略参数即可实现对机器人动作的即时、反馈驱动式纠正。ORPA在预训练控制策略基础上添加了轻量型、反馈条件化模块,该模块直接在关节空间预测残差调整量,使系统可在运行时调整自身行为。我们在ALOHA平台上的一组对精度敏感的操控任务上评估ORPA,结果显示,与基线控制策略和基于规则的逆运动学纠正方法相比,ORPA在成功率和从微小扰动中恢复的能力方面均有提升。
英文摘要
Robotic manipulation policies trained via imitation learning, such as Action Chunking with Transformers (ACT), can achieve strong performance under ideal conditions but often remain sensitive to small execution errors and distribution shifts. Correcting these failures typically requires dataset aggregation and full-policy retraining, which is computationally expensive and unsuitable for real-time deployment. In this work, we propose Online Residual Policy Adaptation (ORPA), a framework that enables immediate, feedback-driven correction of robot actions without modifying the underlying policy parameters. ORPA augments a pretrained control policy with a lightweight, feedback-conditioned module that predicts residual adjustments directly in joint space, allowing the system to adapt its behavior at runtime. We evaluate ORPA on a set of precision-sensitive manipulation tasks using the ALOHA platform, demonstrating improvements in success rate and recovery from small perturbations compared to baseline control policies and rule-based inverse kinematics corrections.