arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

字典序多目标在策略蒸馏

Lexicographic Multi-Objective On-Policy Distillation

Doseok Jang, Jon Ander Campos, Youran Qi

arXiv 2610.02359首次发表:更新:

发表机构

Cohere; Mila, Université de Montréal(Cohere; Mila,蒙特利尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出字典序多目标在策略蒸馏(LMOPD),通过门控选择缺陷专家并投影修正,在多个数学基准上保留高优先级能力,优于现有基线。

AI 中文摘要

从可验证奖励中进行的强化学习(RLVR)通常优化答案的正确性,然而有用的语言模型行为还需要高质量的推理和简洁的响应。现有的多奖励后训练方法通常对奖励进行标量化或组合专家,而没有明确保护奖励的优先级顺序。当权衡不对称时,这会带来问题:例如,简洁性不应以正确性为代价来提高。我们引入了字典序多目标在策略蒸馏(LMOPD),这是一种多教师方法,用于在明确优先级下整合奖励特化策略。对于每次学生轨迹,LMOPD选择第一个目标中门控检测到缺陷的专家,然后局部投影其中心化的对数策略修正,以移除与更高优先级专家相抵触的组成部分。我们在三个数学基准上,以两个和四个专家设置评估了30B-A3B混合专家Transformer模型,并测量保留的专家增益。在两位专家设置下,LMOPD的点估计完全保留了准确性和推理质量的增益,同时获得了简洁性增益的46.9%。在四位专家设置下,它保留了准确性和推理正确性增益的约90%,而评估的次优基线仅保留了约57%。匹配的四专家消融实验表明,字典序路由优于随机路由,并且投影进一步增强了最高优先级能力。在这两种规模下,LMOPD比我们评估的现有基线更有效地保留了最高优先级能力,展示了明确优先级在专家整合中的价值。

英文摘要

Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\%$ of the conciseness gain. With four experts, it retains $\approx90\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\approx57\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.

Comments24 pages, 3 figures, 5 tables; includes appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑