发表机构
Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型对齐依赖英语偏好数据导致其他语言性能欠佳的问题,提出CRPO框架,通过层级结构联合优化语言内与跨语言偏好,在多语言实验中性能优于标准方法。
AI 中文摘要
大语言模型的对齐高度依赖以英语为中心的高质量偏好数据,这往往导致其在其他语言中的性能欠佳。本文提出跨语言排序偏好优化(Cross-Lingual Ranking Preference Optimization, CRPO),这一新颖框架利用英语中的稳健偏好知识来促进目标语言的偏好对齐。我们在目标语言与英语的平行偏好对中设计了一种层级结构,以联合优化语言内和跨语言偏好,从而增强语言适应性和输出质量。CRPO 基于 LambdaLoss 框架构建,超越了基于二元比较的优化,提供了多个候选响应之间的相对排序信号。我们在五种具有不同资源规模的语言上开展实验,结果表明,CRPO 在遵循指令和知识利用能力方面始终优于标准方法。值得注意的是,在各种加权方案下观察到的稳健性能提升,进一步验证了我们的层级设计在多语言设置中的经验有效性。此外,我们的研究结果表明,CRPO 显著提高了奖励边际和理想响应的对数概率,为跨语言对齐构建了更稳定的偏好流形。
英文摘要
The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-lingual Ranking Preference Optimization~(CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language. We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra- and inter-lingual preferences, thereby enhancing language adaptation and output quality. Building on the LambdaLoss framework, CRPO goes beyond the binary comparison based optimization by providing a relative ranking signal across multiple candidate responses. Our experiments across five languages with varying resource scales demonstrate that CRPO consistently outperforms standard approaches in both instruction-following and knowledge utilization capability. Notably, the robust performance gains observed across various weighting schemes further validate the empirical effectiveness of our hierarchical design in a multilingual setup. Furthermore, our findings highlight that CRPO significantly improves both reward margins and the log-probability of desirable responses, contributing to a more stable preference manifold for cross-lingual alignment. Our code is available at https://github.com/dltmddbs100/CRPO.
CommentsEMNLP 2026 Main