arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LayerRoute:通过LoRA保持质量的适应性层跳过方法实现高效LLM推理

LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference

Prateek Kumar Sikdar

arXiv 2609.13682首次发表:更新:

发表机构

Accenture(埃森哲)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LayerRoute通过逐层硬门控路由与LoRA联合微调,实现自适应层跳过,在Qwen2.5-0.5B上保持质量并加速1.02-1.06倍。

AI 中文摘要

我们提出了LayerRoute,一种用于自适应Transformer层跳过的参数高效方法,该方法将逐层硬门控路由(通过直通估计器训练)与联合LoRA微调相结合。LayerRoute为Qwen2.5-0.5B-Instruct中的24个Transformer块中的每一个增加了一个轻量级逐层路由器(约21.5K参数)和LoRA适配器(秩为8,约1.08M参数),并在门控正则化的语言建模目标下联合训练两者。在10个独立种子的训练运行中,LayerRoute在每次运行中都收敛到相同的跳过模式结构——在所有10个种子中,一组一致的9个中间层(8-16层)变得可跳过——并且在每次运行中都实现了真实、经过验证的墙钟加速(1.02倍至1.06倍,平均1.04倍)。在测试的每种配置中,质量均得到保持或提升:联合LoRA适配在所有10个种子中相对于未修改的主干网络实现了困惑度改进(在两个评估分割上的平均差分别为-1.16和-1.11)。我们进一步验证了路由器执行了真实的、非平凡的逐输入计算:在可跳过层中的门控决策改变了87%至100%的保留样本的实际跳过/运行结果,确认了真正的输入依赖路由,而非固定的剪枝模式。LayerRoute在单个A100上训练时间不到7分钟,并且除了路由决策本身之外,增加了可忽略的开销。我们报告了完整的可复现性方法,包括对决定路由器逐输入决策因素的系统性诊断调查,作为本工作的一部分。

英文摘要

We introduce LayerRoute, a parameter-efficient method for adaptive transformer layer-skipping that combines per-layer hard-gated routing (trained via a straight-through estimator) with joint LoRA fine-tuning. LayerRoute augments each of the 24 transformer blocks in Qwen2.5-0.5B-Instruct with a lightweight per-layer router (~21.5K parameters) and LoRA adapters (rank 8, ~1.08M parameters), training both jointly under a gate-regularized language-modeling objective. Across 10 independently-seeded training runs, LayerRoute converges to an identical skip-pattern structure in every run - a consistent set of 9 middle layers (8-16) becomes skip-eligible in all 10 seeds - and delivers genuine, verified wallclock speedup in every run (1.02x-1.06x, mean 1.04x). Quality is preserved or improved in every configuration tested: joint LoRA adaptation yields a perplexity improvement over the unmodified backbone in all 10 seeds (mean delta = -1.16 and -1.11 across the two evaluation splits used). We further verify the router performs genuine, non-trivial per-input computation: gate decisions in skip-eligible layers change the actual skip/run outcome for 87-100% of held-out samples, confirming real input-dependent routing rather than a fixed pruning pattern. LayerRoute trains in under 7 minutes on a single A100 and adds negligible overhead beyond the routing decision itself. We report our full reproducibility methodology, including a systematic diagnostic investigation into what determines the router's per-input decisions, as part of this work.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑