arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

机器人学习中的进度奖励建模:全面综述

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

Jianshu Zhang, Keliang Wu, Haoran Lu, Anbang Liu, Ce Zhang, Weijie Yin, Chengxuan Qian, Xiyuan Yang, Zhenyu Pan, Guo Ye, Han Liu

arXiv 2607.21655首次发表:更新:

发表机构

Northwestern University; Carnegie Mellon University; University of Wisconsin–Madison; University of California, Santa Barbara; University of Illinois Urbana-Champaign(西北大学; 卡内基梅隆大学; 威斯康星大学麦迪逊分校; 加利福尼亚大学圣塔芭芭拉分校; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该综述针对机器人学习中进度奖励建模缺乏共享框架的问题,分三步组织领域研究,包括研究进度模型接口、构建信号方法及支持数据基准,将模型相关要点联系起来,还总结了局限并探讨未来方向。

AI 中文摘要

机器人学习发生在具有大行为空间的动态环境中。终端成功信号仅告知机器人任务是否完成,无法说明当前行为是在取得进展、保持不变还是在撤销早期进展。因此,近期研究越来越多地探索在任务执行期间提供反馈的进度奖励。然而,当前文献缺乏共享框架,现有方法使用不同的观测、目标规范、输出信号、监督源和评估协议,难以比较和理解其结果实际验证了什么。在本综述中,我们对机器人学习的进度奖励建模提供统一观点。我们分三个相关步骤组织该领域。首先研究进度模型的接口,从外部定义问题,即模型接收什么信息以及产生何种形式的进度信号。接着深入模型内部,研究用于构建此信号的方法,揭示进度估计和奖励生成背后的不同假设和机制。最后检查支持这些方法的数据和基准,展示如何获得进度监督以及不同评估实际衡量的内容。这三个视角将进度模型是什么、如何构建以及其质量如何验证联系起来。我们进一步总结了当前方法的主要局限性并讨论了未来研究方向。

英文摘要

Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations, goal specifications, output signals, supervision sources, and evaluation protocols. This makes it difficult to compare them and understand what their results actually validate. In this survey, we provide a unified view of progress reward modeling for robotic learning. We organize the field in three connected steps. We first study the interface of a progress model. This defines the problem from the outside by asking what information the model receives and what form of progress signal it produces. We then move inside the model and study the methods used to construct this signal. This reveals the different assumptions and mechanisms behind progress estimation and reward generation. Finally, we examine the data and benchmarks that support these methods. This shows how progress supervision is obtained and what different evaluations actually measure. Together, these three perspectives connect what a progress model is, how it is built, and how its quality is validated. We further summarize the main limitations of current approaches and discuss future research directions.

CommentsProject page: https://github.com/sterzhang/Awesome-Progress-Models

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑