基于偏好的对齐中的分组失真保证
Groupwise Distortion Guarantees for Preference-Based Alignment
浏览论文内容
中文总结 AI 辅助
针对偏好对齐中的社会福利失真问题,提出GLHF算法,利用每用户一次比较学习分组条件策略,在每组上渐近达到最优失真界,实验显示其显著降低各组及最差组失真。
中文摘要 AI 辅助
基于偏好的对齐方法,如基于人类反馈的强化学习(RLHF)和基于人类反馈的纳什学习(NLHF),通过聚合成对偏好来学习大语言模型(LLM)策略,但一个自然的目标是最大化社会福利(平均基数效用),而仅凭比较无法确定这一点。Gölz、Haghtalab 和 Yang(GHY)通过失真来衡量这一差距:最佳固定彩票(对响应的分布)的福利与所学彩票的福利之间的最坏情况比率。他们证明,当每个用户获得相同的彩票时,NLHF 是最优的。然而,基于账户的 LLM 拥有其用户的信息,可以为不同的人提供不同的彩票。我们提出了一种高效算法 GLHF,它从每个用户的一次比较中学习一个单一的分组条件策略。在个体 Bradley–Terry 比较下,GLHF 在预先指定的、可能重叠的集合中的每个分组上同时渐近匹配 GHY 的最优总体失真界,其样本复杂度随分组数量的对数增长,并与最小分组规模成反比。对于偏好相似的分组,一个更精确的保证是,当成员共享一个可行的偏好响应时,失真趋近于一。在使用人类咖啡评分和合成 LLM 生成评分的实验中,GLHF 降低了每个评估分组的失真,并相对于 NLHF 和其他分组无关的基线大幅降低了最差分组的失真。
英文摘要
Preference-based alignment methods such as reinforcement learning from human feedback (RLHF) and Nash learning from human feedback (NLHF) aggregate pairwise preferences to learn an LLM policy, but a natural goal is maximizing social welfare (average cardinal utility), which comparisons alone need not identify. Gölz, Haghtalab, and Yang (GHY) measure the gap by distortion: the worst-case ratio between the welfare of the best fixed lottery (distribution over responses) and of the learned lottery. They show NLHF is optimal when every user receives the same lottery. Account-based LLMs, however, have information about their users and can serve different lotteries to different people. We give an efficient algorithm, GLHF, that learns a single group-conditioned policy from one comparison per user. Under individual Bradley--Terry comparisons, GLHF asymptotically matches GHY's optimal population distortion bound simultaneously on every group in a prespecified, possibly overlapping collection, with sample complexity growing logarithmically in the number of groups and inversely with the smallest group mass. A sharper guarantee for groups with similar preferences approaches distortion of one when members share a feasible favorite response. In experiments using human coffee ratings and synthetic LLM-generated ratings, GLHF lowers distortion in every evaluated group and substantially reduces worst-group distortion relative to NLHF and other group-agnostic baselines.
发表机构
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。