arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12578cs.AI

从协作到能力:将路由LLM专家内化为紧凑推理器

From Collaboration to Capability: Internalizing Routed LLM Experts into Compact Reasoners

发表机构山东大学 · 浙江大学
查看机构详情
  • Shandong University(山东大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Frank Nie, Shuyao Wang, Ethan B. Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出Rivet框架,通过专家增强强化学习和轨迹内化,将路由LLM专家的协作能力内化到紧凑推理器中,在竞赛数学基准上显著提升准确率并泛化至科学推理。

中文摘要 AI 辅助

一个紧凑的控制器可以通过选择咨询对象、制定请求并整合其响应来协调更强大的专家。我们研究从控制器的决策以及专家的推理和代码中学习,是否能在移除专家后改善其生成能力。我们引入\textsc{Rivet}用于\emph{协作内化}:专家增强的强化学习将共享的结果信号应用于控制器决策和返回的专家片段,并通过格式感知的监督训练整合完整的成功交互,实现经过验证的轨迹内化。部署后的控制器生成推理、代码和交互结构,并支持本地Python执行,无需外部LLM。在七个竞赛数学基准上,RIVET-1.7B和RIVET-4B的平均准确率分别达到$28.25\\%$和$44.16\\%$;第二阶段在移除专家后将RIVET-4B的准确率提升了$6.49$个百分点,GPQA-Diamond的结果提供了向科学推理泛化的证据。消融实验显示,普通轨迹监督和额外格式加权带来了收益,支持了对经过验证的协作内容和结构进行训练的有效性。

英文摘要

A compact controller can coordinate stronger experts by selecting whom to consult, formulating requests, and integrating their responses. We study whether learning from both the controller's decisions and the experts' reasoning and code improves its generation after expert removal. We introduce \textsc{Rivet} for \emph{collaboration internalization}: expert-augmented reinforcement learning applies a shared outcome signal to controller decisions and returned expert spans, and verified trajectory internalization consolidates complete successful interactions through format-aware supervised training. The deployed controller generates reasoning, code, and interaction structure with local Python execution and no external LLM. Across seven competition-mathematics benchmarks, RIVET-1.7B and RIVET-4B achieve average accuracies of $28.25\%$ and $44.16\%$; Stage~II improves RIVET-4B's accuracy after expert removal by $6.49$ points, and GPQA-Diamond results provide evidence of generalization to scientific reasoning. Ablations show gains from ordinary trajectory supervision and additional format weighting, supporting the effectiveness of training on the content and structure of verified collaborations.

↑