arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分歧解码:无训练的能力融合

Divergence Decoding: Training-Free Capability Fusion

Yimi Wang, Hao Li, Shuo Yang, He Cao, Dechen Zhang, Ziang Wu, Zhiyuan Yan, Fanyang Mo, Li Yuan

arXiv 2607.27248首次发表:更新:

发表机构

Peking University; International Digital Economy Academy (IDEA); The University of Hong Kong(北京大学; 国际数字经济学院; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对通用大模型缺领域知识、专家模型逻辑弱的问题,提出无训练的Divergence Decoding框架,通过Jensen-Shannon散度自适应路由,在科学基准上融合两类模型能力,性能优于单模型基线。

AI 中文摘要

尽管大型语言模型在推理方面表现出色,但这些通用模型往往缺乏专业科学领域的知识。相反,领域模型(专家模型)虽然知识丰富,却存在专业化副作用,包括逻辑能力下降等问题。为解决这一困境,我们提出了Divergence Decoding,一种用于能力融合的无训练框架。它将推测解码的“草稿-验证”框架重构为自适应路由机制,核心是使用Jensen-Shannon散度监测两个模型在每个token处的分布分歧。当专家模型表现出显著分歧时,该方法将其识别为潜在推理风险,并即时将控制权路由至通用模型,从而动态注入通用推理能力,同时保留领域专业知识,实现通用模型与专家模型的推理时策略组合。我们在Qwen和Llama系列等不同模型家族的GPQA、ChemBench、ChemCoTBench等具有挑战性的科学基准上评估了Divergence Decoding。实验结果表明,Divergence Decoding的性能优于领域专用模型和通用模型,有效超越了大多数单模型基线的性能。这表明Divergence Decoding提供了一种通用的无训练范式,可通过自适应推理时协作融合不同大型语言模型的能力。

英文摘要

While large language models excel in reasoning, these generalists often lack knowledge for specialized scientific domains. Conversely, domain models~(specialists), while knowledgeable, suffer from specialization side-effects including diminished logic and reduced robustness.To address this dilemma, we introduce Divergence Decoding, a training-free framework for capability fusion. It reconstructs the "draft-and-verify" skeleton of speculative decoding into an adaptive routing mechanism. The core is using Jensen-Shannon divergence to monitor the distributional disagreement between the two models at each token. When the specialist exhibits significant divergence, our method identifies it as a potential reasoning risk and instantaneously routes control to the generalist. This allows the dynamic injection of general reasoning while preserving domain expertise, achieving inference-time policy composition of the generalist and the specialist.We evaluate Divergence Decoding across diverse model families (Qwen and Llama series) on challenging scientific benchmarks (GPQA, ChemBench, and ChemCoTBench). Experimental results demonstrate that Divergence Decoding outperforms both the domain-specialized and general-purpose models, effectively surpassing the performance of most single-model baseline. This suggests that Divergence Decoding provides a general, training-free paradigm for fusing diverse LLM capabilities through adaptive inference-time collaboration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑