arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

梯度对齐的配对选择用于个性化偏好优化

Gradient-Aligned Pair Selection for Personalized Preference Optimization

Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou

arXiv 2610.00061首次发表:更新:

发表机构

Kent State University; EvenUp, USA; Auburn University(肯特州立大学; EvenUp(美国); 奥本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出GAP-DPO,通过梯度对齐的配对选择优化个性化偏好,提升风格保真度与生成质量。

AI 中文摘要

个性化大型语言模型(LLMs)要求将生成行为与用户特定偏好对齐,而非仅仅追求整体质量。虽然直接偏好优化(DPO)为偏好学习提供了稳定的框架,但其在个性化设置中的有效性关键取决于偏好配对的选择方式。现有方法通常依赖启发式标准,如基于似然的极端值,这些方法将优化与明确的用户效用脱钩,可能导致个性化效果退化。我们通过分析期望用户效用的梯度与DPO更新方向之间的一阶交互,将个性化偏好学习形式化为一个几何对齐的优化问题。我们的分析揭示,在离策略采样下,当偏好边际与效用梯度方向对齐时,DPO更新从纯粹的错误纠正信号转变为类似强化学习的更新。这一视角将配对选择暴露为一个几何决策,它决定了偏好优化是推进还是阻碍个性化。基于这一洞见,我们提出了GAP-DPO(几何对齐偏好DPO),一种迭代算法,它执行效用感知的、几何对齐的配对选择,并通过逐轮再生成来控制分布偏移。在个性化文本生成基准上的实验表明,与标准DPO变体相比,GAP-DPO持续提高了风格保真度、偏好对齐和生成质量。综合来看,我们的结果确立了梯度对齐作为个性化偏好优化的统一原则,并证明配对选择是优化几何的内在组成部分,而非启发式预处理步骤。

英文摘要

Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optimization from explicit user utility and can lead to degraded personalization. We formalize personalized preference learning as a geometry-aligned optimization problem by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. Our analysis reveals that, under off-policy sampling, the DPO update transitions from a purely error-corrective signal to a reinforcement-like update when preference margins are directionally aligned with utility gradients. This perspective exposes pair selection as a geometric decision that governs whether preference optimization advances or hinders personalization. Motivated by this insight, we propose GAP-DPO (Geometry-Aligned Preference DPO), an iterative algorithm that performs utility-aware, geometry-aligned pair selection while controlling distribution shift via epoch-wise regeneration. Experiments on personalized text generation benchmarks show that GAP-DPO consistently improves stylistic fidelity, preference alignment, and generation quality compared to standard DPO variants. Together, our results establish gradient alignment as a unifying principle for personalized preference optimization and demonstrate that pair selection is an intrinsic component of the optimization geometry rather than a heuristic preprocessing step.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑