arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09796cs.LG

噪声偏好标签下无元数据的元重加权直接偏好优化

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

Hua Qu, Yifan Li, Xiaodong Yuan

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对DPO性能依赖偏好数据质量问题,提出双层优化框架、无任务元知识驱动方法及结合中心差分近似与LoRA微调的可扩展训练方案,经实验验证该方法能在不同噪声率下提升训练性能。

中文摘要 AI 辅助

直接偏好优化(DPO)已成为使大语言模型(LLMs)与人类偏好对齐的重要方法,因其无需显式奖励建模和强化学习优化。但其性能严重依赖偏好数据质量,现实中噪声偏好数据会削弱对齐性能。为此提出双层优化框架,在一定假设下可恢复干净数据下的DPO最优解。还推导了非对称标签翻转噪声下可学习加权函数的先验形式。考虑到高质量元数据难获取,提出无任务元知识驱动方法,即便无元数据也能元学习。结合中心差分近似与LoRA微调降低LLM元学习中高阶梯度的高成本,开发可扩展训练方案。在TL;DR摘要和Anthropic HH单轮对话实验表明,该方法在不同噪声率下比多个DPO基线提高了训练性能。

英文摘要

Direct Preference Optimization (DPO) has become an important method for aligning large language models (LLMs) with human preferences because it removes the need for explicit reward modeling and reinforcement learning. However, its performance depends heavily on the quality of preference data, and noisy preference data in real-world settings can weaken alignment performance. To address this issue, we propose a bilevel optimization framework and prove, under some idealized conditions, that this framework can recover the DPO optimum under clean data. We further derive a prior form for the learnable weighting function under label-flipping noise. Considering that high-quality metadata may be difficult to obtain, we propose a prompt augmentation consistency method that enables meta-learning even when metadata is completely unavailable. To reduce the high cost of higher-order gradients in LLM meta-learning, we combine central-difference approximation with LoRA fine-tuning and develop a scalable training scheme. Experiments on TL;DR summarization and Anthropic Helpful and Harmless dialogue show that the proposed method improves alignment performance over multiple DPO baselines under different noise rates.

发表机构

  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑