arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31360cs.AI

通过皮层区域的视角看大语言模型中的层程序

Programs-of-Layers in LLMs through the Lens of Cortical Areas

Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers

首次发表
浏览论文内容

中文总结 AI 辅助

本研究复现并深入分析PoLar系统,证明动态路由层程序优于固定前向传播,但学习路由器失效,且错误纠正程序脆弱,揭示其与丘脑-皮层协调的相似性。

中文摘要 AI 辅助

大语言模型(LLM)的推理传统上是一个固定深度、固定顺序的前向传播过程,逐层执行,无论输入难度如何。人脑并非如此运作:人脑以丘脑作为中央枢纽,根据需求将信息灵活地路由到皮层的所有区域。Li等人(2026)最近展示了一种他们称之为“层程序”(PoLar)的系统,表明如果变压器的层被视为函数库而非固定序列,则可以赋予它们类似的灵活性。当每个输入通过跳过或重复的连续层块的适应性序列进行动态路由时,性能优于标准前向传播。我们比原始论文更详细地重建了PoLar的诊断性蒙特卡洛树搜索(MCTS),并将其应用于5个模型。我们重现了PoLar的若干发现:跳过优于标准传播,重复优于跳过,而两者结合则优于任一单独使用。较短的程序足以解决较简单的问题,而较难的问题则需要更多的层重复。然而,我们未能复现其关于单次推理的学习路由器的核心主张:其排名第一的预测始终退化回标准传播,尽管其排名前k的预测程序整体上确实显示出真实的准确率提升。除复现外,我们发现少量通用程序足以解决大多数问题。我们还对这些程序的结构和鲁棒性进行了更深入的分析:例如,我们发现纠正错误的程序高度脆弱:即使撤销程序中的单个编辑,通常也会破坏纠正效果。将此与大脑的路由机制联系起来,PoLar反映了丘脑-皮层协调原则,类似于皮层区域般的变压器层。我们在此https URL公开发布代码。

英文摘要

Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all regions of the cortex according to demand. Li et al. (2026) recently showed, with a system they call program-of-layers (PoLar), that transformers can be given an analogous flexibility if their layers are treated as a library of functions rather than a fixed sequence. Performance improves over the standard forward pass when each input is dynamically routed through an adaptive sequence of skipped or repeated contiguous layer blocks. We reconstructed PoLar's diagnostic MCTS in more detail than the original paper and applied it across 5 models. We reproduced several of PoLar's findings: skipping outperformed the standard pass, repeating outperformed skipping, and combining both outperformed either alone. Shorter programs sufficed for easier questions, while harder questions required more layer repeats. However, we failed to replicate the main claim regarding their learned router for single-shot inference: its top-ranked prediction consistently collapsed back to the standard pass, even though its top-k predicted programs, taken together, did show a real accuracy gain. Beyond reproduction, we find that a small number of generic programs are enough to solve most of the questions. We also provide a much deeper analysis of these programs' structure and robustness: for example, we found that programs that correct errors are highly brittle: undoing even a single edit inside a program typically breaks the correction. Connecting this to the brain's routing mechanisms, PoLar mirrors principles of thalamo-cortical coordination between cortical-area-like transformer layers. We publicly release the code at https://datexis.github.io/RE-PoLar/

发表机构

  • Berliner Hochschule für Technik (BHT)(柏林工业大学)
  • Humboldt-Universität zu Berlin(柏林洪堡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑