arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24900cs.LG

逆强化学习通过模仿人类帮助对齐人工智能

Inverse RL Helps Align AI by Imitating Humans

Michał Wiliński, Liu Leqi, Chirag Nagpal

首次发表
浏览论文内容

中文总结 AI 辅助

研究语言模型对齐问题,受逆强化学习启发提出PARED方法,通过恢复示范的隐式奖励来改进基础策略,无需特定任务偏好注释,经实验验证其有效性及可用于上下文对齐。

中文摘要 AI 辅助

语言模型对齐旨在使模型行为可靠地反映诸如有用性、安全性和遵循指令等理想属性。当前方法通常使用示范上的监督微调或基于验证器或人类反馈得出奖励的强化学习。这些范式未充分探索一个重要问题:仅示范能否产生可检查、重用和在线优化以对齐人工智能的隐式奖励?受逆强化学习启发,我们引入了从示范估计的投影对齐奖励(PARED)。PARED将专家示范背后的隐式奖励恢复为一小组响应级特征上的显式函数,由轻量级判别器学习,该判别器在特征空间中将示范与策略自身样本分开。与标准奖励模型不同,PARED不需要特定任务的偏好注释:示范提供特定任务监督,可通过人工智能反馈作为额外监督维度增强。通过涉及推理时重新排序和对抗性在线强化学习的实验,我们表明恢复的奖励可在无监督损失的情况下改进基础策略,并且在标准监督微调后进行优化时会带来进一步提升。此外,我们证明PARED可用于上下文对齐,即单个策略可针对不同受众的偏好进行定制。

英文摘要

Language model alignment aims to make model behavior reliably reflect desirable properties such as helpfulness, safety, and instruction following. Current approaches typically use supervised fine-tuning on demonstrations or reinforcement learning with rewards derived from verifiers or human feedback. These paradigms leave an important question underexplored: can demonstrations alone yield an implicit reward that can be inspected, reused, and optimized on-policy to align AI? Motivated by inverse reinforcement learning, we introduce Projected Alignment Reward Estimated from Demonstrations (PARED). PARED recovers the implicit reward underlying expert demonstrations as an explicit function over a small set of response-level features, learned by a lightweight discriminator that separates demonstrations from the policy's own samples in this feature space. Unlike a standard reward model, PARED requires no task-specific preference annotations: demonstrations provide the task-specific supervision, which can be augmented with AI feedback as additional dimensions of supervision. Through experiments involving inference-time reranking and adversarial on-policy RL, we show that the recovered reward improves a base policy without a supervised loss and yields further gains when optimized after standard supervised fine-tuning. Additionally, we demonstrate that PARED can be used for contextual alignment, in which a single policy can be tailored to the preferences of different audiences.

↑