arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33838cs.LGcs.CL

ChemOPD:面向多任务化学推理的多教师在线策略蒸馏

ChemOPD: Multi-Teacher On-Policy Distillation for Multi-Task Chemical Reasoning

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Xuemin Chen, Tianshu Yu

AI总结:

ChemOPD通过亲和性引导的专家分组和锚定残差多教师在线策略蒸馏,在ChemCoTBench上提升了多任务化学推理能力,证明专业化和能力整合是相关但独立的设计问题。

AI中文摘要:

大型语言模型日益被期望在统一模型中支持多种化学推理能力。一种方法是分别开发专门能力,并通过多教师在线策略蒸馏加以整合,但这引出了两个问题:专业化应如何组织,以及专家指导应如何整合?我们提出了ChemOPD,它解决了这两个问题。我们从监督微调梯度中估计任务亲和性,并求解一个约束混合整数规划(MIP)来构建部分重叠的专家小组。在蒸馏过程中,我们保留一个在所有任务上训练的通才教师,使得专家指导补充而非取代其监督。我们的锚定残差目标逐步增加路由专家在学生生成响应上的贡献。在ChemCoTBench上,亲和性引导的专业化相对于通才教师产生了任务相关的收益,并在若干能力上超越了语义任务分组。然而,更强的教师端性能并不自动产生更强的学生:在相同的专家和路由下,锚定残差OPD在大多数报告的指标上优于仅专家蒸馏,并实现了可用教师收益的更大份额。这些结果凸显了专业化和能力整合是化学推理中相互关联但截然不同的设计问题。

英文摘要:

Large language models are increasingly expected to support diverse chemical reasoning capabilities within a unified model. One approach is to develop specialized capabilities separately and consolidate them through multi-teacher on-policy distillation, but this raises two questions: how should specialization be organized, and how should specialist guidance be integrated? We introduce ChemOPD, which addresses both. We estimate task affinities from supervised fine-tuning gradients and solve a constrained mixed-integer program(MIP) to construct partially overlapping specialist groups. During distillation, we retain a generalist teacher trained on all tasks so that specialist guidance supplements rather than replaces its supervision. Our anchor-residual objective gradually increases the routed specialist's contribution on student-generated responses. On ChemCoTBench, affinity-guided specialization produces task-dependent gains over the generalist teacher and improves several capabilities beyond semantic task grouping. Yet stronger teacher-side performance does not automatically yield stronger students: with the same specialists and routes, anchor-residual OPD improves most reported metrics over specialist-only distillation and realizes a larger share of the available teacher gains. These results highlight specialization and capability integration as connected but distinct design problems in chemical reasoning.

↑