内在机器人奖励:重用VLA表示进行自主评估与策略改进
Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement
浏览论文内容
中文总结 AI 辅助
IRR方法重用VLA系统的视觉表示和演示终点,通过参考库和评分操作实现自主结果评估与策略改进,降低集成成本并减少人工评分。
中文摘要 AI 辅助
视觉-语言-动作(VLA)系统已经汇集了机器人学习的两项宝贵资源:丰富的视觉表示和成功任务执行的演示。内在机器人奖励(IRR)提出将这些资源用于第二个互补目的:评估机器人自身的结果并为策略改进提供反馈。成功的演示终点定义了任务特定的参考,而策略的冻结视觉编码器提供了评估新结果的特征空间。核心奖励机制在现有流程中增加了一个参考库和一个评分操作,无需单独的学习评估器或额外的感知骨干网络。我们的立场是,这种重用为降低集成工作量、高效奖励计算和减少重复的人工结果评分提供了一条有前景的途径。基于视觉奖励和从经验中学习的已有研究,IRR将这些想法引入机器人现有的感知和演示流程中。一个可操作的COMAU Racer 3演示器已达到技术就绪级别4(TRL 4)。这一实验室基础支持下一步研究:将内部结果评估与物理策略改进联系起来。我们提出了奖励公式、核心研究问题以及一种将奖励可靠性与任务成功和监督工作量联系起来的评估方法。预期贡献是一种可重用的方法,用于从工业机器人系统中已有的数据和经验中学习和改进。
英文摘要
Vision-language-action (VLA) systems already bring together two valuable resources for robot learning: rich visual representations and demonstrations of successful task execution. Intrinsic Robot Rewarding (IRR) proposes to use these resources for a second, complementary purpose: evaluating the robot's own outcomes and providing feedback for policy improvement. Successful demonstration endpoints define task-specific references, and the policy's frozen visual encoder provides the feature space in which new outcomes are assessed. The core reward mechanism adds a reference bank and a scoring operation to the existing pipeline, without requiring a separate learned evaluator or an additional perception backbone. Our position is that this reuse offers a promising route to lower integration effort, efficient reward computation, and reduced recurring human outcome scoring. Building on established research in visual rewards and learning from experience, IRR brings these ideas into the robot's existing perception and demonstration pipeline. An operational COMAU Racer 3 demonstrator is available at technology readiness level 4 (TRL 4). This laboratory foundation supports the next research step: connecting internal outcome evaluation to physical policy improvement. We present the reward formulation, central research questions, and an evaluation methodology linking reward reliability to task success and supervision effort. The intended contribution is a reusable approach to learn and improve from the data and experience already available in industrial robot systems.
发表机构
- Deggendorf Institute of Technology(德根多夫理工学院)
机构由 AI 辅助整理,请以论文原文为准。