arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25200cs.LGcs.AIcs.CL

学习用于多目标对齐的Plackett-Luce模型混合

MoPLEx: Estimating Plackett-Luce Mixture Models for Multi-Objective Alignment

Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对异质偏好下多向排序的模型混合问题,提出MoPLEx算法,通过扩充排序并结合梯度估计与期望最大化,在偏好优化任务上显著优于基线方法。

中文摘要 AI 辅助

我们考虑在给定标注者的多向排序响应(这些响应可能代表异质的潜在偏好)的情况下,学习k个Plackett-Luce模型混合的问题。该问题在AI对齐和偏好优化中有诸多应用。现有研究已从成对比较的角度探讨了Bradley-Terry模型的混合。然而,当k超过排序长度m的一半时,混合模型在理论上是不可识别的。我们提出一种高效实现方案以解决该限制,该方案首先通过从基础语言模型生成新响应,将排序扩充至更大规模,随后在输入嵌入空间中进行基于梯度的估计以降低推理成本。基于此过程,我们设计了包含上述两个步骤的期望最大化算法来拟合Plackett-Luce模型的混合,该算法名为MoPLEx。我们开展了大量实验以验证该方法:首先,在参数规模达340亿的模型上,我们证明基于梯度的近似方法估计真实概率的误差低于5%;其次,在偏好优化数据集上,与使用单一排序的基线方法及Bradley-Terry模型混合的基线方法相比,MoPLEx将聚类准确率平均提升43.7%,排序准确率平均提升15.2%。这些结果表明,MoPLEx通过测量梯度间的对齐关系,在处理来自异质偏好的多向排序时具有有效性。

英文摘要

We study learning a mixture of $k$ Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons. However, estimating a mixture of multi-way ranking models can become theoretically unidentifiable when $k$ exceeds $m/2$, where $m$ is the ranking length. We design an efficient algorithm to address this issue by first augmenting the rankings to a larger size (e.g., generating comparisons from a base model), followed by a gradient-based estimation to reduce inference cost (in the input embedding space). With this procedure in mind, we then fit a mixture of Plackett-Luce (PL) models via an expectation-maximization-style iteration, or MoPLEx in short. We conduct extensive experiments to verify this algorithm. First, we find that the gradient-based approximation estimates true probabilities with less than 5% error on models with up to 34 billion parameters. Second, MoPLEx improves clustering and ranking accuracy by an average of 43.7% and 15.2% over baselines using a single PL model or a mixture of Bradley-Terry models, on UltraFeedback and PERSONA datasets. These results demonstrate the effectiveness of MoPLEx for tackling multi-way rankings following heterogeneous preferences through measuring alignment via gradients.

发表机构

  • Northeastern University(东北大学)
  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑