arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19144cs.CLcs.AIcs.LG

大语言模型偏好对齐的零阶范式

A Zeroth-Order Paradigm for LLM Preference Alignment

Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出零阶对齐方法ComPO,利用比较oracle提取偏好方向信息,避免似然位移,并在多种LLM上实现更优的长度控制胜率。

中文摘要 AI 辅助

直接偏好对齐方法因其计算和内存效率而被广泛用于使大型语言模型(LLMs)与人类偏好对齐。然而,似然位移促使人们寻找从具有小似然边际的偏好对中提取信息的替代方法。在本文中,我们提出并分析了一种基于比较或acles的零阶对齐方法——基于比较的偏好优化(ComPO)。ComPO从这些偏好对中提取方向信息,而无需直接优化可微的偏好损失。我们为其基本的离线方案在平滑性、梯度稀疏性以及oracle与潜在目标之间的兼容性条件下建立了收敛保证。我们进一步引入了在线ComPO,它保留了离线比较机制,并使用未标记的策略生成来相对于参考策略进行反向KL控制。遵循偏好微调的覆盖视角,我们在局部覆盖和分布内成对奖励准确性的条件下,为基本约束方案建立了性能保证。在Mistral、Llama、Gemma-2、Qwen3和Gemma-3模型上的实验表明,与现有直接对齐方法相比,包括长度控制的胜率在内的性能有所提升,成对级别的诊断提供了与缓解似然位移一致的证据。

英文摘要

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we propose and analyze Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles. ComPO extracts directional information from these pairs without directly optimizing a differentiable preference loss on them. We establish a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective. We further introduce online ComPO, which retains the offline comparison mechanism and uses unlabeled policy generations for reverse-KL control relative to a reference policy. Following the coverage perspective of preference fine-tuning, we establish a performance guarantee for a basic constrained scheme under local coverage and in-distribution pairwise reward accuracy. Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 models demonstrate improvements over existing direct alignment methods, including length-controlled win rates, with pair-level diagnostics providing evidence consistent with mitigating likelihood displacement.

发表机构

  • University of California, Berkeley(加州大学伯克利分校)
  • New York University(纽约大学)
  • DAMO Academy, Alibaba Group U.S.(阿里巴巴集团美国达摩院)
  • Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑