arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

T-Router:通过参数高效强化学习学习丘脑路由进行推理

T-Router: Learning Thalamic Routing for Reasoning with Parameter-Efficient Reinforcement Learning

Liuxian Ma, Jiale Dai, Jiaqi Li, Lu Mi

arXiv 2609.39109首次发表:更新:

发表机构

College of Artificial Intelligence, Tsinghua University; State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University; Beijing Institute for General Artificial Intelligence(清华大学人工智能学院; 北京大学智能科学与技术学院通用人工智能国家重点实验室; 北京通用人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

T-Router通过可寻址存储库和循环控制器实现计算复用,在8.95B模型上仅用0.466%参数提升推理,超越全参数GRPO和LoRA,平均AIME准确率从48.33升至60.56。

AI 中文摘要

参数高效强化学习旨在通过一个紧凑的可训练接口与预训练模型交互,从而提升推理能力。我们引入了丘脑路由器(T-Router),它将适应性集中在已完成计算的复用上。一个压缩的、可寻址的存储库保存块级变化;一个深度循环控制器调节其选择及相对规模的写回。这种耦合赋予了丘脑上下文相关路由一个具体的计算形式:学习接收层使用哪些早期贡献,以及以何种影响力使用。正确性奖励训练该接口,同时保持骨干参数和层顺序不变。在8.95B参数的骨干模型上,T-Router分配了41.73M参数(占骨干的0.466%),并在GSM8K强化学习后达到了83.64±1.16的MathAvg,而全参数GRPO在三次评估轮次中为73.79±1.83。在可比的参数预算和匹配的重试次数下,它超过了LoRA的77.28±1.95 MathAvg,改善了所有三个任务族,并将平均AIME准确率从48.33提升至60.56。容量控制的比较有利于可寻址的块变化和循环上下文;独立的搜索训练将接口扩展到工具介导的推理。这些结果确立了受控计算复用作为参数高效推理强化学习有效途径的地位。

英文摘要

Parameter-efficient reinforcement learning aims to improve reasoning with a compact trainable interface to a pretrained model. We introduce the Thalamic Router (T-Router), which concentrates adaptation on the reuse of completed computations. A compressed, addressable bank preserves block changes; a depth-recurrent controller conditions their selection and relative-scale writeback. This coupling gives thalamic context-dependent routing a concrete computational form: learn which earlier contributions a receiving layer uses, and with what influence. Correctness rewards train the interface while preserving backbone parameters and layer order. On an 8.95B-parameter backbone, T-Router allocates 41.73M parameters (0.466% of the backbone) and achieves 83.64 +/- 1.16 MathAvg after GSM8K RL, compared with 73.79 +/- 1.83 for full-parameter GRPO across three evaluation rounds. At a comparable parameter budget and with matched retries, it exceeds LoRA's 77.28 +/- 1.95 MathAvg, improving all three task families and raising mean AIME accuracy from 48.33 to 60.56. Capacity-controlled comparisons favor addressable block changes and recurrent context; separate search training extends the interface to tool-mediated reasoning. These results establish controlled computation reuse as an effective route to parameter-efficient reasoning reinforcement learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑