arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34484cs.ROcs.LG

ARS:用于机器人学习的智能体奖励系统

ARS: Agentic Reward System for Robot Learning

  • AIRC, Midea Group(美的集团人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Sheng Hu, Weiyi Lu, Lingbing Zeng, Gan Weng, Weiwei Zhang, Kai Xie, Xiaofeng Mou, Yi Xu

AI总结:

本文提出智能体奖励系统(ARS),利用通用视觉语言模型通过自适应视觉检查和子智能体验证,改进机器人进度奖励建模,抑制虚假进度,并在仿真和真实机器人上验证其有效性。

AI中文摘要:

进度奖励建模是估计机器人行为如何随时间改变任务进度的问题。可靠的估计需要区分有意义的状态变化与失败尝试及与任务无关的动作。我们引入了智能体奖励系统(ARS),一个使用通用视觉语言模型(VLM)进行进度奖励建模的推理框架,无需额外的奖励模型训练。给定一个离线轨迹和任务指令,ARS使用自适应视觉检查进行事件提议和验证。一个子智能体提议一个与任务相关的事件时间线,主智能体在估计每帧进度之前对其进行验证和修订。ARS可以整合可选的终端结果标签和视觉参考来指导其判断。它还可以审计来自外部奖励模型的进度估计。我们在一个受控的语义不匹配基准测试中,使用27B VLM评估ARS,并在仿真和真实机器人上进行下游策略学习。基准测试揭示,即使在简单的抓取和放置场景中,几个评估的奖励基线也会将虚假进度分配给错误对象的操作。ARS更好地抑制了这些错误,并在仿真策略学习中优于这些基线。我们进一步证明,ARS支持从混合质量的离线经验中进行长视界策略学习,在工业洗衣机装配线的全尺寸实验室复制品上进行真实机器人多螺丝紧固。这些结果表明,结构化推理和验证可以提高通用VLM在机器人奖励建模中的实用性。代码位于此https URL。

英文摘要:

Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training. Given an offline trajectory and a task instruction, ARS uses adaptive visual inspection for both event proposal and verification. A subagent proposes a task-relevant event timeline, which a primary agent verifies and revises before estimating per-frame progress. ARS can incorporate optional terminal outcome labels and visual references to inform its judgments. It can also audit progress estimates from external reward models. We evaluate ARS with a 27B VLM on a controlled semantic-mismatch benchmark and downstream policy learning in simulation and on a real robot. The benchmark reveals that several evaluated reward baselines assign spurious progress to wrong-object manipulation even in simple pick-and-place scenes. ARS better suppresses these errors and outperforms these baselines in simulation policy learning. We further demonstrate that ARS supports long-horizon policy learning from mixed-quality offline experience on real-robot multi-screw fastening in a full-scale laboratory replica of an industrial washing-machine assembly line. These results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot reward modeling. Code is at https://github.com/midea-ai/ars

↑