发表机构
Sogang University; Korea University(西江大学; 高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现指令质量是偏好学习的隐性瓶颈,提出基于奖励信号与评分标准引导的LLM反馈的指令优化流程,可提升偏好数据质量,增强LLM对齐效果。
AI 中文摘要
偏好学习利用响应对优化模型,而这些响应对的信息量从根本上由生成它们的指令决定。我们发现指令质量是偏好学习中的一个隐性瓶颈:低质量或模糊的指令会限制响应质量的分布,削弱优质的被选响应,进而弱化偏好信号。通过最优N和最差N分析,我们表明指令质量会约束采样响应质量的上限和下限。基于这一观察,我们提出一种指令优化流程,该流程利用奖励信号筛选弱指令,并通过基于 rubric(评分标准)的大语言模型(LLM)反馈对其进行修正,在不丢弃样本的情况下提升偏好数据质量。在离线和在线偏好学习场景下,对多个模型和基准的实验显示,与原始数据及其他数据优化策略相比,该方法实现了广泛的对齐效果提升。进一步分析表明,指令优化可提升可达到的响应质量,并与以响应为中心的偏好数据整理形成互补。总体而言,指令质量是决定大语言模型(LLM)对齐所需偏好信号信息量的关键因素。代码可在该 https URL 获取。
英文摘要
Preference learning optimizes models using response pairs, yet the informativeness of these pairs is fundamentally shaped by the instructions from which they are generated. We identify instruction quality as a hidden bottleneck in preference learning: low-quality or ambiguous instructions restrict the response-quality distribution, limiting strong chosen responses and weakening preference signals. Through Best- and Worst-of-N analyses, we show that instruction quality constrains both the ceiling and floor of sampled response quality. Motivated by this observation, we introduce an instruction-refinement pipeline that selects weak instructions using reward signals and revises them with rubric-guided LLM feedback, improving preference data without discarding examples. Across offline and online preference learning settings, experiments on multiple models and benchmarks show broad alignment improvements over original data and alternative data-improvement strategies. Further analyses indicate that instruction refinement raises achievable response quality and complements response-centric preference data curation. Overall, instruction quality emerges as a key factor governing how informative preference signals are formed for LLM alignment. Code is available at: https://github.com/01choco/instruction-refinement/
CommentsPreprint