arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

能量正则化模仿学习用于力与功感知的机器人操作

Energy-Regularized Imitation Learning for Force- and Work-Aware Robotic Manipulation

Toshiki Otani, Hiromu Taketsugu, Norimichi Ukita

arXiv 2609.18164首次发表:更新:

发表机构

Toyota Technological Institute(丰田工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出能量正则化模仿学习方法,通过可微分能量预测器将机械功作为正则项微调操作策略,在RLBench上降低2.1%平均功并提升成功率,实现无需显式动力学模型的节能操作。

AI 中文摘要

本文研究能量感知操作作为一个物理基础的学习问题。我们利用关节力矩和角位移定义了一个关节空间机械功代理,并训练了一个可微分的能量预测器,该预测器从机器人状态和动作中估计此功。该预测器将模拟器侧不可微分的物理量转化为可微分的正则化项,用于微调预训练的操作策略。我们在RLBench上使用RVT-2实例化了该框架,并评估了12个涉及物体接触、关节运动、放置、推和清扫的操作任务。所提出的微调将平均机械功从208.8J降低到204.4J(即降低2.1%),同时平均任务成功率也从86.2%略微提升至86.9%。这些结果表明,功感知的策略优化可以抑制物理上低效的运动,而无需显式的可微分动力学模型。

英文摘要

This paper studies energy-aware manipulation as a physically grounded learning problem. We define a joint-space mechanical-work proxy from joint torque and angular displacement, and train a differentiable energy predictor that estimates this work from robot states and actions. The predictor converts a non-differentiable simulator-side physical quantity into a differentiable regularizer for fine-tuning a pretrained manipulation policy. We instantiate the framework with RVT-2 on RLBench and evaluate 12 manipulation tasks involving object contact, articulated motion, placement, pushing, and sweeping. The proposed fine-tuning reduces the average mechanical work from 208.8J to 204.4J (i.e., 2.1% reduction), while the mean task success rate also increases slightly from 86.2% to 86.9%. These results show that work-aware policy optimization can suppress physically inefficient motion without requiring an explicit differentiable dynamics model.

CommentsECCV 2026 Workshop on Force-Grounded, Cross-View Articulated Manipulation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑