arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18770cs.LGcs.RO

欲行远,需同行:多样化偏好引出奖励优化的课程设置

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

Taehyung Kim, Jongeun Choi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对AI对齐中服务不足用户的问题,提出CurriPO方法,通过构建树状课程设置适配多样化用户目标,在个性化连续控制任务上使群体满意度达最强基线的1.2-2.1倍且缩短训练时间。

中文摘要 AI 辅助

从人类反馈中学习奖励模型并据此优化策略,是使AI系统与单个用户对齐的一种方法。从公平角度看,现有研究通过开发数据高效且准确的奖励模型,在数据稀缺时仍能捕捉少数群体偏好,以此改进这种对齐效果。我们将这一研究方向推进了一步,提出数据高效且准确的单用户奖励模型还不够:那些在策略层面难以优化其奖励模型的用户,会成为新的服务不足群体。我们从一个观察出发:某个用户的奖励模型从初始策略优化起来可能很容易,而另一个用户的则不然。我们认为,在用户群体足够多样化的情况下,容易优化和难以优化的奖励模型之间会自然形成一种课程设置。基于这一见解,我们提出了CurriPO,它会构建一个树状结构的课程设置以适配多样化的用户特定目标,仅需一次遍历即可覆盖整个群体。具体而言,CurriPO会自动在多样化的用户奖励模型上构建课程设置,使其能从现有课程设置分支,并复用之前纳入课程设置的奖励模型。据我们所知,这是首个明确利用多用户结构解决AI对齐中优化问题的研究。在模拟环境下的个性化连续控制任务上开展的大量实验表明,CurriPO实现了最强基线1.2至2.1倍的群体满意度,同时大幅减少了训练时间。进一步分析显示,该改进很大程度上源于传统优化被忽视的服务不足用户。

英文摘要

Learning a reward model from human feedback and optimizing a policy against it is one approach to aligning AI systems with individual users. From a fairness perspective, existing work improves such alignment by developing data-efficient and accurate reward models that capture minority preferences despite scarce data. We push this line of inquiry one step further and argue that data-efficient and accurate per-user reward models are not sufficient: users whose reward models are difficult to \textit{optimize} at the policy level can become a new underserved group. We start from the observation that one user's reward model can be easy to optimize from the initial policy while another's is not. We argue that, given a sufficiently diverse user population, a curriculum naturally emerges between easy- and hard-to-optimize reward models. Building on this insight, we propose CurriPO, which grows a tree-structured curriculum to accommodate diverse user-specific objectives, covering the population in a single traversal. Specifically, CurriPO automatically constructs a curriculum over diverse user reward models, allowing it to branch from the existing curriculum and reuse reward models previously incorporated into the curriculum. To the best of our knowledge, this is the first work to explicitly exploit multi-user structure to address optimization in AI alignment. Extensive experiments on personalized continuous control in a simulated environment show that CurriPO achieves $1.2$--$2.1\times$ the population satisfaction of the strongest baseline while substantially reducing training time. Additional analysis attributes much of this improvement to the users left underserved by conventional optimization.

补充信息

↑