发表机构
Georgia Institute of Technology; Zillow Group(佐治亚理工学院; 齐洛集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多目标对齐的不平衡问题,提出MINT方法,通过按最弱目标排名候选替代加权和排名,在情感支持等任务中显著提升双目标表现并降低不平衡性。
AI 中文摘要
将语言智能体同时对齐多个目标是基于偏好的训练中持续存在的失效模式:当目标以加法方式组合时,优化会崩溃到最容易改进的那个目标,而牺牲其余目标,因此支持智能体学会表现得热情却不提供实际帮助。根本问题在于加法奖励缺乏平衡的概念。我们引入MINT(MIN-selection preference disTillation,最小选择偏好蒸馏),这是对偏好蒸馏的一行修改:我们不再按奖励的加权和对采样候选进行排名,而是按其最弱目标进行排名,在DPO目标保持不变的情况下,将最平衡的候选蒸馏为最优,将最不平衡的候选蒸馏为最差。这是从加法到最坏情况选择的广义均值族中p→负无穷的极限。在合作情感支持和对抗性谈判任务中,最小选择同时提升了两个目标,同时大幅降低了它们的不平衡性;在情感支持任务中,它将较弱的维度从0.37提升到0.64(p<10^-40),超过了人类专家,并在完整多轮生成中保持效果。逐轮分析得出我们的核心发现:最小选择纠正不平衡的程度与参考策略的不平衡程度成正比,其益处会在交互中持续,持续时间恰好与该不平衡的持续时间一致。
英文摘要
Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no real help. The root issue is that an additive reward has no notion of balance. We introduce Mint (MIN-selection preference disTillation), a one-line change to preference distillation: rather than ranking sampled candidates by a weighted sum of rewards, we rank them by their weakest objective, distilling the best-balanced candidate over the most lopsided one with an unchanged DPO objective. This is the p -> negative infinity limit of a generalized-mean family spanning additive to worst-case selection. Across cooperative emotional support and adversarial negotiation, min-selection lifts both objectives while sharply cutting their imbalance; on emotional support it raises the weaker axis from 0.37 to 0.64 (p < 10^-40), surpassing human experts and persisting across full multi-turn rollouts. A turn-by-turn analysis yields our central finding: min-selection corrects imbalance in proportion to how imbalanced the reference policy is, and its benefit endures over an interaction precisely as long as that imbalance does.