arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12302cs.LG

奖励函数设计框架:从目标到特征再到人类对齐的奖励函数

A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions

Di Yang Shi, W. Bradley Knox

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出一套三步式奖励函数设计框架,首次将奖励项选择简化为因果DAG上的最小成本部分覆盖问题,能保持无冲突可行权重区域,为非专家设计人类对齐的奖励函数提供了可行方案。

中文摘要 AI 辅助

我们提出了一个正式流程,使非专家能够实例化并迭代人类对齐的奖励函数,即符合给定轨迹偏好排序的奖励函数。给定用自然语言描述的任务,我们的流程分三步生成线性奖励函数:第一步,将任务的目标提炼为一组基本目标,并推导能捕获这些基本目标的可测量结果变量;第二步,选择因果代表性的结果变量子集作为奖励项;第三步,通过偏好征询为这些奖励项拟合权重。我们的贡献包括将第一步描述为推导结果变量的引导式工作流,将第二步形式化为将奖励项选择简化为因果有向无环图(DAG)上的最小成本部分覆盖问题,通过最大流在多项式时间内求解,将第三步形式化为将权重拟合视为凸可行性问题,通过现有分离神经由偏好查询迭代收窄求解。据我们所知,这是首个能保持确定性无冲突可行权重区域的奖励设计方法,该区域可通过具有O(n log κ)次偏好查询的分离神经由期望容差收窄。

英文摘要

We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. reward functions that adhere to a given preference ordering over trajectories. Given a task described in natural language, our process produces a linear reward function in three steps: distill the task's objectives into a set of fundamental objectives and derive measurable outcome variables that capture those fundamental objectives, select a causally representative subset of outcome variables as the reward terms, and fit weights to those reward terms via preference elicitation. Our contributions describe the first step and formalize the latter two steps. The first is a guided workflow for deriving outcome variables. The second is a reduction of reward term selection to minimum-cost partial cover on a causal DAG, solved in polynomial time via max-flow. The third is a geometric framing of weight fitting as a convex feasibility problem iteratively narrowed by preference queries, solved by existing separation oracle methods. To the best of our knowledge, this is the first reward-design method that maintains a deterministically conflict-free feasible weight region, narrowed to a desired tolerance via a separation oracle with O(n log κ) preference queries.

补充信息

↑