arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30597cs.LGcs.CL

PLC-DPO:带噪声与歧义偏好优化中的后验标签修正

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

发表机构韩国科学技术院人工智能学院 · 韩国世宗国立大学(原韩国科学技术研究院世宗校区)
查看机构详情
  • KAIST AI(韩国科学技术院人工智能学院)
  • KENTECH(韩国世宗国立大学(原韩国科学技术研究院世宗校区))

机构由 AI 辅助整理,请以论文原文为准。

Boryeong Cho, Sumyeong Ahn, Se-Young Yun

首次发表
浏览论文内容

中文总结 AI 辅助

PLC-DPO针对DPO假设偏好全可靠的缺陷,提出后验标签修正方法,经57个数据集-模型-基准单元验证,其平均胜率优于DPO及次优方法,且在各类测试中表现稳定。

中文摘要 AI 辅助

直接偏好优化(DPO)通过成对比较简化了对齐过程,但假设所有观测到的偏好均可靠。真实数据常违背这一假设,产生反向、弱或歧义标签,导致有害的策略更新。为解决该问题,我们提出后验标签修正DPO(PLC-DPO),通过将每对的训练信号归类为干净、翻转或平局情况,实现对偏好的鲁棒优化。核心思路是使用校准后的策略-参考边际作为在线证据,采取适当的修正动作。这将噪声偏好学习重新定义为主动修正监督方向与强度,而非仅过滤可疑样本。在57个数据集-模型-基准单元上,PLC-DPO相对于DPO取得了最佳平均胜率(60.5,次优方法为55.5)。注入噪声和平局压力测试、人类分歧分析及自确认诊断进一步表明,该归类机制保持稳定,可区分翻转对与弱定向对。

英文摘要

Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.

补充信息

↑