Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
超越成对:通过排名选择建模增强大语言模型对齐
机构 * Institute of Operations Research and Analytics(运营研究与分析研究所) ; National University of Singapore(新加坡国立大学) ; Department of Analytics and Operations(分析与运营系) ; NUS Business School(新加坡国立大学商学院)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI、cs.LG
AI总结 本文提出RCPO框架,通过最大似然估计结合排名选择建模,提升大语言模型对齐效果,实验显示其在多种模型和设置中均优于基线方法。
Comments Accepted by The Fourteenth International Conference on Learning Representations (ICLR 2026)