发表机构
AntGroup(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多教师在线策略蒸馏中行为杠杆不平衡问题,提出Soft Clamp方法,通过校准令牌级散度,减少模型过度调用工具情况,在保持决策准确率的同时,降低工具调用循环和重复调用,提升多教师OPD效果。
AI 中文摘要
智能语言模型必须学会何时调用工具、何时使用工具响应以及何时直接回答。这使得多教师在线策略蒸馏成为一种自然的训练策略。我们表明这种策略会引发一种仅从总体损失中无法察觉的行为转变。在双教师工具使用场景中,普通的广义知识蒸馏提高了工具调用召回率,但也使模型倾向于过度调用。我们转而分析行为杠杆不平衡,提出了Soft Clamp方法。在APIGen - MT上,Soft Clamp相对于普通GKD将过度调用从13.7%降至9.0%,同时匹配其决策准确率。在BFCL多轮诊断中,它还减少了GKD变体中的工具调用循环和重复调用。这些结果表明多教师OPD应监测教师信号作用的位置,而非仅关注其总体大小。
英文摘要
Top-$K$ teacher logits make on-policy distillation tractable, but retained teacher mass does not certify student-relative gradient fidelity. We study routed, two-teacher tool-use distillation. In a frozen Qwen3.5-9B audit, the tool teacher ranks the entry token first on 500 tool prompts, and top-32 reinforces it in every matched pair. The response teacher's top-32 instead retains 99.99\% mass yet contains the same token on only 0.4\% of 500 response prompts; student-aware support restores the coordinate in all matched pairs and nearly matches full-vocabulary descent. On the tool route, top-32 preserves the reinforcing direction but substantially attenuates its full-vocabulary magnitude despite mass displayed as 1.000000. Matched restoration connects the response-side omission to behavior: restoring the entry coordinate at every supervised response position lowers full-generation over-calling from $14.2 \pm 2.1\%$ to $3.7 \pm 0.5\%$ across three seeds, but also lowers call recall by 12.4 points. A non-tool placebo changes over-calling by only 0.95 points. A synthetic output-logit-gradient-matched reroute also fails to reproduce exact restoration. We compare three optimization layers: student-aware support changes distilled coordinates, loss shaping changes retained-signal strength, and validation-tuned entry bias shifts inference without retraining. They expose distinct trade-offs in correction scope, system access, and required-call retention. Llama-3.1-8B reproduces the directional support asymmetry under its native JSON protocol. The results causally implicate decision-critical support omission as one contributor in the primary Qwen setting and show why both support fidelity and deployment costs must be audited. Code and aggregate artifacts are available at https://github.com/shen-jiabin/decision-support-opd.
Comments33 pages, 5 figures. Code and aggregate artifacts: https://github.com/shen-jiabin/decision-support-opd