DIAG:面向数据高效数学偏好蒸馏的诊断式迭代对齐与生成
DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation
- Tsinghua University(清华大学)
- Kyoto University(京都大学)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DIAG是一种自适应调整训练分布的框架,通过诊断偏好对产出与生成针对性训练数据,提升数学推理任务中偏好监督的信息价值,在同等训练预算下增强模型性能。
AI中文摘要:
迭代偏好优化是将大语言模型对齐至数学推理任务的核心环节,但其效率常受限于信号稀缺性:随着模型性能提升,静态问题集与模型不断演进的能力愈发不匹配,生成的输出要么过易要么过难,因此不具备信息价值,进而导致有效偏好对的稀缺。我们提出DIAG(Diagnostic Iterative Alignment and Generation,诊断式迭代对齐与生成)框架,该框架自适应调整训练分布,以增加有价值的监督信号,并将训练聚焦于学生模型当前的能力边界附近。DIAG包含两个阶段:(1)诊断有效偏好对产出,通过经验贝叶斯收缩估计器校准探索-利用权衡并分配主题配额,从而优先选择高产出概念;(2)生成针对性训练数据,其中教师模型从学生模型的失败轨迹中合成变体。我们进一步从理论视角解释DIAG是教师介导的、将训练分布向学生能力边界进行KL正则化重加权的近似,在此过程中有效偏好对产出被最大化。实验表明,DIAG在各迭代中提升了产出,且在同等有效训练预算下实现了更强的推理性能,证明其可为数学推理蒸馏出更具信息价值的偏好监督信号。
英文摘要:
Iterative preference optimization is essential for aligning Large Language Models on mathematical reasoning tasks, yet its efficiency is often throttled by signal scarcity: as the model improves, static problem sets become increasingly mismatched to the model's evolving competence, producing rollouts that are either too easy or too hard and therefore non-informative, which leads to a scarcity of valid preference pairs. We propose DIAG, a Diagnostic Iterative Alignment and Generation framework that adaptively reshapes the practice distribution to increase informative supervision and focus training near the student's current competence boundary. DIAG consists of two phases: (1) diagnosing valid preference-pair yield to calibrate the exploration-exploitation trade-off and allocate topic quotas via an Empirical Bayes shrinkage estimator, thereby prioritizing high-yield concepts; and (2) generating targeted practice, where a teacher synthesizes variants from the student's failure traces. We further provide a theoretical view interpreting DIAG as a teacher-mediated approximation to KL-regularized reweighting of the practice distribution toward the student's competence boundary, where valid preference-pair yield is maximized. Experiments show that DIAG boosts yield across iterations and delivers stronger reasoning performance under an iso-effective training budget, demonstrating that it can distill more informative preference supervision for mathematical reasoning.