发表机构
The University of Hong Kong; Yale University; Purdue University(香港大学; 耶鲁大学; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对列表偏好排序标签不确定性,提出点态全变差鲁棒Plackett-Luce目标,将内层最大化简化为排序,在离线和在线设置中均具有理论保证,实验表明能保持干净标签性能并提升噪声鲁棒性。
AI 中文摘要
现有的用于语言模型对齐的鲁棒偏好优化主要研究成对监督,并将鲁棒性置于数据集、提示或偏好对级别。我们转而研究在排序标签不确定性下的列表偏好优化:给定一个提示和一个候选列表,由于标注者不一致、近似平局、有损排序反馈或奖励模型噪声,观察到的列表排序可能是不明确的。我们提出一个点态全变差鲁棒Plackett-Luce目标,该目标直接对候选列表条件下的排序标签进行鲁棒化。鲁棒损失可以精确分解为名义PL损失加上一个最坏情况PL修正,并且最坏情况排序通过将当前隐式分数按升序排序得到,将内层最大化从$K!$枚举减少到$O(K\log K)$。这种易处理的结构产生了强大的离线和在线优化保证。在离线固定列表设置中,鲁棒目标是凸的,投影随机次梯度方法以$O(\epsilon^{-2})$的样本复杂度达到全局$\epsilon$-次优。在在线策略诱导设置中,其中候选列表由当前策略生成,我们建立了弱凸性和$\widetilde O(\epsilon^{-2})$的Moreau包络平稳性。在离线LLM对齐实验中,所提出的鲁棒修正方法在干净标签下基本保持性能,并在噪声下提高了鲁棒性。在线对齐中,它使奖励模型排序的候选扩展更可靠,并改善了奖励模型和外部GPT-4评估指标。
英文摘要
Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We instead study listwise preference optimization under ranking-label uncertainty: given a prompt and a candidate list, the observed ranking over that list may be ambiguous due to annotator inconsistency, near-ties, lossy rankwise feedback, or reward-model noise. We propose a pointwise total-variation robust Plackett--Luce objective that directly robustifies the ranking label conditional on the candidate list. The robust loss admits an exact decomposition into the nominal PL loss plus a worst-case PL correction, and the worst-case ranking is obtained by sorting current implicit scores in ascending order, reducing the inner maximization from $K!$ enumeration to $O(K\log K)$. This tractable structure yields strong offline and online optimization guarantees. In the offline fixed-list setting, the robust objective is convex and projected stochastic subgradient reaches global $ε$-suboptimality with $O(ε^{-2})$ sample complexity. In the online policy-induced setting, where candidate lists are generated by the current policy, we establish weak convexity and $\widetilde O(ε^{-2})$ Moreau-envelope stationarity. Experiments in offline LLM alignment show that the proposed robust correction largely preserves performance under clean labels and improves robustness under noise. In online alignment, it makes reward-model-ranked candidate expansion more reliable and improves both reward-model and external GPT-4 judge metrics.