通过不确定性校准的MOPD在领域专业化过程中保留通用能力
Recovering General Capabilities via Uncertainty-Calibrated Multi-Teacher On-Policy Distillation
查看机构详情
- Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对大语言模型领域专业化时通用能力下降的问题,提出不确定性校准的MOPD方法,在角色扮演和医疗领域实验中提升通用能力且保持垂直领域性能。
中文摘要 AI 辅助
将大语言模型专业化到垂直领域可提升领域特定行为,但常导致推理、编码、指令遵循、创意写作等通用能力下降。我们研究多教师在线蒸馏(Multi-Teacher On-Policy Distillation,MOPD)中的领域-通用权衡,其中专业化学生由领域教师和通用教师在其自身采样的轨迹上进行监督。标准MOPD面临两个局限:普通在线采样很少暴露具有较大教师-学生优势的标记,且仅优势符号无法确定所得更新方向是否可靠。我们提出不确定性校准的MOPD以解决这些局限:双温度采样扩大候选轨迹池,正优势密度过滤选择具有更强正学习信号的轨迹,中心对数似然(Centered Log-Likelihood,CLL)过滤随后计算熵校准的教师认可分数,并根据方向-认可一致性概率性保留标记更新。在角色扮演和医疗领域专业化上的实验表明,我们的方法分别比标准MOPD提升通用能力平均值4.73%和10.84%,同时保持垂直领域性能。消融实验和诊断分析进一步证实,增益并非仅来自更大的部署预算,且所提出的轨迹级和标记级机制解决了其预期的失败模式。
英文摘要
Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities. We study this trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized model learns from domain and general teachers on its own sampled trajectories. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, and advantage sign alone does not establish whether the proposed update direction is reliable. We propose Uncertainty-Calibrated MOPD (UCMOPD), which addresses these limitations through two complementary mechanisms. Golden-Gain Enhancement combines higher-temperature exploration with a standard-temperature anchor and retains trajectories whose positive learning signal matches or exceeds the prompt-specific anchor. Teacher-Endorsement Filtering then uses centered log-likelihood (CLL) to estimate each retained token's plausibility relative to the teacher's uncertainty and probabilistically preserves updates whose directions are supported by that endorsement. Across role-playing and medical-domain specialization, UCMOPD improves the general-capability average over standard MOPD by $4.48\%$ and $7.86\%$, respectively, while maintaining vertical-domain performance. Component ablations and diagnostic analyses support the intended roles of the two mechanisms: exposing and selecting stronger positive signals at the trajectory level and validating update directions through teacher endorsement at the token level.