内在路由器:从冻结的LLM中激发原生技能路由
The Router Within: Eliciting Native Skill Routing from a Frozen LLM
浏览论文内容
中文总结 AI 辅助
本文提出Gavel方法,利用冻结LLM内部状态通过两个线性映射实现无需上下文的技能路由,在Qwen3-32B上超越外部参数方法,最高提升21.9个百分点,并验证了路由能力随骨干增强而提升。
中文摘要 AI 辅助
技能使LLM智能体超越其参数化知识,而技能带来的增益取决于选择正确的技能。部署的框架通过将每个技能的元数据预加载到上下文中进行路由,这会分散智能体的注意力并限制库的大小。检索流程将选择移出上下文,但也移出了智能体的能力范围。我们证明,冻结的智能体LLM在其自身的前向传播中已经携带了路由信号,并且两个线性映射足以在上下文中无技能文本的情况下读出该信号。Gavel(从冻结LLM中一瞥并裁决)分两步读取该信号。一瞥将任务和每个技能的中间层状态通过这两个映射(唯一训练的参数)投影,并根据安装时一次前向传播构建的紧凑每技能库对完整库进行评分。随后,裁决恢复入围技能的向前传播,并读取模型自身的似然和是/否判断,与一瞥作为专家乘积融合。Gavel仅训练一次,即可零样本迁移到三个公共基准和我们的新基准SkillTraj(包含372条模拟智能体轨迹)。在Qwen3-32B上,它优于渐进式披露和检索-重排流程(这些流程增加了1.2B到16B的外部参数),在书面任务上最高提升13.4个百分点,在技能需求出现在中途时最高提升21.9个百分点。路由准确性随骨干网络的提升而提高,在bash智能体框架中,相同的32B模型在Skill-Use上触发正确技能的频率高于在Codex中运行的远为更大的前沿模型。
英文摘要
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library size. Retrieval pipelines move the selection out of the context, but also out of the agent's capability. We show that the frozen agent LLM already carries the routing signal in its own forward passes, and that two linear maps suffice to read it out with no skill text in the context. Our Gavel (Glance And Verdict from a frozen LLM) reads it in two steps. A glance scores the full library by matching the task's mid-layer states against a compact bank that one forward pass builds for each skill at installation, with the two maps as the only trained parameters. A verdict then resumes each shortlisted skill's forward pass, reads the model's own likelihood and yes/no judgment, and fuses both with the glance as a product of experts. Trained once, Gavel transfers zero-shot to three public benchmarks and SkillTraj, our new benchmark of 372 simulated agent trajectories. On Qwen3-32B it outperforms progressive disclosure and retrieve-and-rerank pipelines that add 1.2B to 16B external parameters, by up to 13.4 points on written tasks and up to 21.9 when the need for a skill arises mid-rollout. Routing accuracy improves as the backbone does, and in a bash-agent harness Gavel lets the 32B trigger the right skill on Skill-Use more often than models of up to 1.6T parameters in Codex.
发表机构
- Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。