发表机构
Indian Statistical Institute; Indian Institute of Technology Kanpur(印度统计研究所; 印度理工学院坎普尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对线性奖励对齐中的公理失效问题,提出引入松弛量并优化总松弛与违规次数的方法,在满足PO和PMC公理下,证明最优松弛紧界并实验验证优于线性BTL。
AI 中文摘要
从人类偏好数据中学习是对齐语言模型与人类价值观的主要途径。在线性社会选择中,奖励是提示-响应对的固定特征表示的线性函数,Ge等人[2024]表明,通过最小化任何非递减凸损失(包括BTL)来拟合这样的奖励,无法满足帕累托最优(PO)和多数人一致性(PMC)。此外,当输出要求为线性诱导时,任何仅读取多数关系的规则都无法满足PO。我们询问强制执行这些公理的成本是什么。为此,我们放宽线性模型,允许每个候选者存在松弛量。我们计算满足公理且具有最小总松弛量的松弛线性奖励,并带有边际η,即两个奖励值之间所需的最小差异。我们的解决方案在不对选民或比较收集方式做任何假设的情况下满足公理。当η最多为O(1/m^2)(m为候选者数量)时,我们将最优总松弛量界定为O(1)。此外,我们展示了一个迫使该界限成立的实例,从而得出结论该速率在常数因子内是紧的。在实践中,候选者数量远超特征维度,且只有线性奖励才能在未见过的响应上进行评估。因此,我们引入了一种新方法,对线性部分每次错误比较进行惩罚。它同时最小化总松弛量和违规次数,并通过参数λ在两者之间进行权衡。我们证明总松弛量在λ上是单调但饱和的:提高λ会减少部署的线性奖励的违规次数并增加松弛量,但松弛量保持在(η+Δ√d)⌊m^2/4⌋以下,其中Δ和d分别是特征的直径和维度。在合成和真实偏好数据上的实验证实了我们的理论,并表明我们方法输出的线性奖励优于线性BTL。
英文摘要
Learning from human preference data is the dominant route to aligning language models with human values. In linear social choice, where rewards are linear in a fixed feature representation of prompt-response pairs, Ge et al.[2024] show that fitting such a reward by minimizing any non-decreasing convex loss, including BTL, fails PO and PMC. Moreover, no rule that reads only the majority relation can satisfy PO once the output is required to be linearly induced. We ask what it costs to enforce these axioms anyway. To this end, we relax the linear model to allow per-candidate slack. We compute the relaxed linear reward with the smallest total slack that satisfies the axioms with a margin $η$, the minimum required difference between two reward values. Our solution satisfies the axioms under no assumptions about the voters or how comparisons were collected. We bound the optimal total slack by $O(1)$ when $η$ is at most $O(\frac{1}{m^2})$ for $m$ candidates. Furthermore, we exhibit an instance that forces this bound, concluding that the rate is tight up to constants. In practice, the no. of candidates far exceeds the feature dimension, and only a linear reward can be evaluated on unseen responses. We therefore introduce a new method that charges the linear part for each comparison it gets wrong. It simultaneously minimizes the total slack and the no. of violations, with a parameter $λ$ trading off between them. We show that the total slack is monotone but saturating in $λ$: raising it reduces the violations of the deployed linear reward and increases the slack, yet the slack stays below $(η+Δ\sqrt{d})\lfloor m^2/4\rfloor$, where $Δ$ and $d$ are the diameter and dimension of the features, respectively. Experiments on both synthetic and real-life preference data corroborate our theory and show that the linear reward output by our method beats linear BTL.
Comments22 pages, 5 figures