arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05872cs.CLcs.AI

MACRO:Transformer层的马尔可夫链路由

MACRO: Markov Chain Routing of Transformer Layers

Paweł Batorski, Abtin Pourhadi, Akylgali Aitaza, Przemysław Spurek, Paul Swoboda

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出无需修改底层参数的MACRO框架,将Transformer层路由建模为马尔可夫策略,在多基准上实现优于基线及Dr. LLM的性能,同时大幅缩短路由搜索时间。

中文摘要 AI 辅助

标准大型语言模型(LLM)按顺序执行层。动态层路由即搜索包含层重复、跳过及其他操作的不同层执行路径,可提升性能。现有路由方法常需更新模型权重、对每个测试实例运行昂贵的搜索循环,或在推理时需要真实标签。本研究提出Transformer层的马尔可夫链路由框架(MACRO),该框架可在不修改底层参数的情况下学习LLM架构的特定任务路由。MACRO将层路由建模为依赖上下文的马尔可夫策略,该策略以层索引、计算预算阶段、方向位移和算子上下文为条件,支持跳过、重复和残差隐状态添加操作。马尔可夫路由分布通过训练数据反馈更新,并使用top-k维特比算法解码以分离高概率候选程序。我们在多个开源权重LLM上的多种推理和知识基准上评估了MACRO。与未路由基线相比,MACRO实现了+5.0%的平均准确率提升,在小型模型上的提升最大。我们以+7.2%的优势优于最佳动态路由方法Dr. LLM,同时将路由搜索时间减少了9.4倍(从14.8小时降至1.6小时)。我们的代码可在该httpsURL公开获取。

英文摘要

Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path through layers involving layer repetitions, skips and other moves, can improve performance. Existing routing approaches often require updating model weights, running expensive search loops per test instance, or demand ground-truth labels during inference. In this work, we propose Markov Chain Routing of Transformer Layers (MACRO), a framework that learns task-specific routes over LLM architectures without modifying underlying parameters. MACRO models layer routing as a context-dependent Markov policy conditioned on layer indices, computation budget phases, directional displacements, and operator context, supporting skip, repeat, and residual hidden-state addition operations. The Markov route distribution is updated via feedback on training data and decoded using a top-k Viterbi algorithm to isolate high-probability candidate programs. We evaluate MACRO across diverse reasoning and knowledge benchmarks on multiple open-weight LLMs. MACRO achieves a +5.0% average accuracy improvement over the unrouted baselines, with largest gains on small models. We outperform the best dynamic routing approach Dr. LLM by +7.2%, while reducing route-search time 9.4x (from 14.8 to 1.6 hours). Our code is publicly available at https://github.com/Batorskq/MACRO.

↑