arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迭代奖励设计的鸟瞰视角

A Bird's-Eye View of Iterative Reward Design

Logan Mondal Bhamidipaty, Lauren Robson, Linda Petrini, Shengrui Lyu, Kamal Ndousse

arXiv 2610.04364首次发表:更新:

发表机构

Anthropic Fellows Program; University of Edinburgh(Anthropic 研究员计划; 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出迭代奖励设计基准(BIRD),统一配置与评估空间,在多个环境中识别简单设计选择以提升性能,并强调简单基线的优势。

AI 中文摘要

在强化学习中设计有效的奖励函数通常需要大量的专业知识和反复试验。近期的工作通过基于大语言模型的系统自动化了这一过程,这些系统利用策略反馈生成并迭代改进奖励代码。然而,这些方法往往难以比较,因为它们在实现细节、反馈假设和评估环境上存在差异。为解决这一问题,我们引入了迭代奖励设计基准(BIRD),该基准在统一的配置和评估空间中表达现有方法。这使我们能够直接比较算法,消融单个设计选择,并在匹配的反馈条件和策略训练预算下原型化新组件。在MuJoCo、Meta-World、Assistax和HumanoidBench上,我们识别出一小组简单的设计选择,这些选择持续提升性能。将这些选择组合起来,其性能显著优于先前工作中评估的方法。我们的结果凸显了简单基线的优势,并激励进一步研究何时额外的算法复杂性会改善迭代奖励设计。代码可在https://this URL获取。

英文摘要

Designing effective reward functions in RL typically requires substantial expertise and trial and error. Recent work automates this process with LLM-based systems that generate and iteratively improve reward code using policy feedback. However, these methods are often hard to compare because they differ in implementation details, feedback assumptions, and evaluation environments. To address this, we introduce a Benchmark for Iterative Reward Design (BIRD) that expresses existing methods in a unified configuration and evaluation space. This lets us compare algorithms directly, ablate individual design choices, and prototype new components under matched feedback conditions and policy-training budgets. Across MuJoCo, Meta-World, Assistax, and HumanoidBench, we identify a small set of simple design choices that consistently improve performance. Combining these choices yields significantly better performance than the evaluated methods from prior work. Our results highlight the strength of simple baselines and motivate further study of when additional algorithmic complexity improves iterative reward design. Code is available at https://github.com/safety-research/bird.

Comments30 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑