arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08189cs.AIcs.CL

动态路由器需要记忆吗?HeRo:面向高效LLM推理的历史感知路由

Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference

Hongjin Lin, Wentao Wan, Keze Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有动态层路由忽视跨层决策依赖的问题,提出历史感知路由(HeRo),通过线性注意力维护显式路由记忆,在冻结骨干上训练轻量组件,显著提升性能保留并降低推理成本。

中文摘要 AI 辅助

动态层路由通过学习为单个令牌跳过层来降低大型语言模型(LLM)的推理成本。然而,现有方法将每个路由决策视为仅以当前隐藏状态为条件的局部操作,这种表述忽略了跨深度路由的序列性和路径依赖性:早期决策影响下游路由器所见的表示,而层使用目标则联合耦合所有决策。我们提出历史感知路由(HeRo),一种动态路由框架,通过引入路由器记忆机制来维护跨模型深度的显式路由状态,从而解决这一不匹配问题。该记忆通过线性注意力构建,将先前的路由分数及其引发的残差更新增量聚合为紧凑的历史表示。在每个被路由的层,路由器联合基于该累积状态和当前隐藏表示来选择要执行的分支。在令牌级前馈网络(FFN)路由的实例化中,HeRo仅在冻结骨干上训练轻量级路由器和适配器,无需修改预训练参数。在Llama 3.1-8B、Llama 2-7B和Llama 2-13B上,HeRo在十个基线中始终实现最高的总体性能保留。在Llama 3.1-8B上,它绕过26.87%的模型参数,同时在七个基准上达到稠密模型性能的100.24%,并在更紧的计算预算下绕过38.82%的模型参数时保留97.01%的性能。消融研究证实,移除路由历史始终降低性能,尤其是在多步推理和代码生成上最为显著,验证了显式路由记忆比仅依赖隐藏状态能实现更准确和自适应的动态路由。

英文摘要

Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers for individual tokens. Existing methods, however, treat each routing decision as a local operation conditioned solely on the current hidden state which is a formulation that overlooks the sequential, path-dependent nature of routing across depth: earlier decisions shape the representations seen by downstream routers, and the layer-usage objective couples all decisions jointly. We propose History-Aware Routing (HeRo), a dynamic routing framework that resolves this mismatch by introducing a router memory mechanism to maintain an explicit routing state across model depth. The memory is constructed via linear attention, incrementally aggregating preceding routing scores and their induced residual updates into a compact history representation. At each routed layer, the router conditions jointly on this accumulated state and the current hidden representation to select the executed branch. Instantiated for token-wise FFN routing, HeRo trains only lightweight routers and adapters on a frozen backbone, requiring no modification to pretrained parameters. Across Llama 3.1-8B, Llama 2-7B, and Llama 2-13B, HeRo consistently achieves the highest aggregate performance retention among ten baselines. On Llama 3.1-8B, it bypasses 26.87% of model parameters while achieving 100.24% of dense model performance across seven benchmarks, and retains 97.01% while bypassing 38.82% of model parameters under a tighter computation budget. Ablation studies confirm that removing routing history consistently degrades performance, most notably on multistep reasoning and code generation, validating that explicit routing memory enables more accurate and adaptive dynamic routing than solely conditioning on hidden state.

发表机构

  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑