arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25136cs.AI

数据更少,对齐更好:用于偏好优化的数据中心多评估器一致性方法

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization

Zhengtao Yao, Runhao Li, Xupeng Chen, Jiayi Cheng, Chenqian Le, Michael Yue, Siheng Wang, Haoyan Xu, Yuqi Li, Chenhao Wei, Zhengdao Li, Rongchao Zhang, Guang Ya… 展开作者

Zhengtao Yao, Runhao Li, Xupeng Chen, Jiayi Cheng, Chenqian Le, Michael Yue, Siheng Wang, Haoyan Xu, Yuqi Li, Chenhao Wei, Zhengdao Li, Rongchao Zhang, Guang Yang, Yidong Wang, Junhao Dong

首次发表
浏览论文内容

中文总结 AI 辅助

研究聚焦偏好优化,提出DMAPO方法,从目标策略生成候选响应,经多评估器评估、校正后保留高一致性示例。实验表明该方法数据效率高,能提升模型在MT - Bench等指标上的表现,且改变评估器或准则对下游性能影响小。

中文摘要 AI 辅助

偏好优化研究通常在固定数据时改变训练目标。本文提出是否一小部分高置信度的策略内响应能提供可靠学习信号的问题。方法DMAPO从目标策略生成候选响应,用专门评估器评估帮助性、事实性和简洁性,应用过程批评校正,仅保留高一致性的可取或不可取示例。该程序在54236个Mistral - 7B候选中接受1871个(3.45%)。基于此训练的KTO在MT - Bench上达到7.50,与text - davinci - 003参考相比长度控制胜率达95.5%,IFEval提示准确率达57.3%。独立成对评估也表明DMAPO优于SimPO。改变评估器模型或准则会改变所选示例,但对下游性能影响不大。另一主干研究接受率为3.41%,性能提升较适度。总体而言,共识过滤为通用指令的偏好优化提供了数据高效途径,代价是额外的策算计算和对评估器判断的依赖。

英文摘要

Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can provide a reliable learning signal. Our method, DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimization), generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only high-consensus desirable or undesirable examples. This procedure accepts 1,871 of 54,236 Mistral-7B candidates (3.45%). KTO trained on this set reaches 7.50 on MT-Bench, 95.5% length-controlled win rate against a text-davinci-003 reference, and 57.3% IFEval prompt accuracy. Independent pairwise evaluation also favors DMAPO over SimPO: GPT-4o yields a net win rate of 23.3 points on 129 held-out prompts and 24.0 points on 200 out-of-distribution LMSYS-Chat prompts; Claude Opus 4.7 yields 24.1 points on the held-out set. Changing the evaluator model or rubric alters the selected examples but has little effect on downstream performance. A second-backbone study yields a similar 3.41% acceptance rate, although its performance gains are more modest. Across these experiments, consensus filtering offers a data-efficient route to preference optimization for general instructions, at the cost of additional curation compute and dependence on evaluator judgments.

发表机构

  • University of Southern California(南加州大学)
  • New York University(纽约大学)
  • Columbia University(哥伦比亚大学)
  • University of California, Berkeley(加利福尼亚大学伯克利分校)
  • City College of New York, CUNY(纽约市立大学城市学院)
  • Stevens Institute of Technology(史蒂文斯理工学院)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑