arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从宪法到控制:用于对齐语言模型的可解释奖励

From Constitutions to Control: Interpretable Rewards for Aligning Language Models

Johann D. Gaebler, Calvin Isley, Max Lamparth, Stephen Casper, Sharad Goel

arXiv 2609.33086首次发表:更新:

AI 中文总结

本研究提出基于评分标准的框架,将宪法转化为可解释、可调的奖励模型,通过重新加权维度实现可控对齐,并缓解标签偏差。

AI 中文摘要

当前对齐语言模型的方法往往难以知道什么行为被奖励,或难以以有针对性的方式改变该奖励。特别是,基于偏好的标准方法将多种考量合并为总体人类判断,掩盖了驱动所得奖励的因素,而基于原则的方法则指定了高层价值观,但未完全将其操作化。为解决这一差距,我们开发了一个基于评分标准的框架,将通用宪法转化为可解释且可调的奖励模型,利用宪法引导的AI反馈来估计各评分标准项的初始权重。然后,我们重新加权这些维度以构建用于训练的修改后奖励。在政治对齐与安全-帮助性权衡的实验中,重新加权单个维度可预测地改变目标行为,且在很大程度上独立进行,同时在对齐目标冲突之间进行权衡。我们表明,同一框架可以通过减少偏好判断中编码的标签偏差(包括谄媚和人口统计偏差)对训练奖励的影响来缓解这些偏差。我们的结果表明,源自宪法的可解释奖励可以将高层对齐原则转化为更透明和可控的模型行为。

英文摘要

Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgments, obscuring what drives the resulting reward, while principle-based methods specify high-level values without fully operationalizing them. To address this gap, we develop a rubric-based framework to transform a general-purpose constitution into an interpretable and tunable reward model, using constitution-guided AI feedback to estimate initial weights for the constituent rubric items. We then reweight those dimensions to construct modified rewards for training. Across experiments on political alignment and safety-helpfulness tradeoffs, reweighting individual dimensions predictably changes targeted behaviors largely independently while navigating tradeoffs between conflicting alignment objectives. We show that the same framework can mitigate label bias encoded in preference judgments -- including sycophancy and demographic bias -- by reducing their influence on the training reward. Our results demonstrate that constitution-derived, interpretable rewards can translate high-level alignment principles into more transparent and controllable model behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑