arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07851cs.LGcs.CL

TEMPER:用于表达性残差路由的张量化高效流形约束参数化

TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

Yuxuan Gu, Wuyang Zhou, Huijun Xing, Danilo Mandic

中文总结 AI 辅助

该研究针对残差路由方法参数随流增长过快的问题,提出TEMPER方法,用张量网络参数化生成器,在语言建模等任务上性能相当或更优,参数效率显著提升。

中文摘要 AI 辅助

残差连接依赖静态残差通路,是训练深度神经网络的关键。超连接(HC)通过整合多条残差流并学习动态信息流,提升残差路由的表达能力;而流形约束超连接(mHC)变体则通过双随机残差混合稳定训练。然而现有方法存在生成器层面的瓶颈:它们使用密集、非结构化的生成器进行分支前聚合、残差混合及分支后重分配,导致参数数量随流数量快速增长。为解决该问题,我们提出用于表达性残差路由的张量化高效流形约束参数化方法(TEMPER),将这些生成器表示为输入-流、特征、输出-流模式上的多路张量,并使用张量网络对其进行参数化。这种结构化低秩公式被证明能保留依赖 token 的流形约束路由接口,同时大幅减少参数增长。它还提升了可解释性与直观性:i)张量秩控制学习到的路由子空间的维度,满秩时可恢复密集路由;ii)生成器近似误差限制了路由 logits 的差异,进而限制了路由块输出的差异。综合实验表明,TEMPER 在语言建模和常识推理任务上与现有方法相当或更优,同时所需额外参数大幅减少。在8条残差流下,TEMPER 取得了最佳 CORE 分数,同时比 mHC 少用约84%的额外参数,展现出更强的性能-参数效率权衡。

英文摘要

Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expressivity of residual routing by incorporating multiple residual streams and learning dynamic information flow, while manifold-constrained (mHC) variants stabilize training through doubly stochastic residual mixing. However, a generator-level bottleneck remains in existing methods: they use dense, unstructured generators for pre-branch aggregation, residual mixing, and post-branch redistribution, which results in parameter count growing rapidly with the number of streams. To address this issue, we propose \underline{\textbf{T}}ensorized \underline{\textbf{E}}fficient \underline{\textbf{M}}anifold-constrained \underline{\textbf{P}}arameterization for \underline{\textbf{E}}xpressive Residual \underline{\textbf{R}}outing (\textbf{TEMPER}), which represents these generators as multi-way tensors over the input-stream, feature, and output-stream modes, and parameterizes them using tensor networks. Such a structured low-rank formulation is shown to preserve token-dependent manifold-constrained routing interface while substantially reducing parameter growth. It also promotes interpretability and intuition, as: i) tensor ranks control the dimensionality of the learned routing subspace, with full ranks recovering dense routing; while ii) the generator approximation errors bound differences in routing logits and, consequently, in the routed-block outputs. Comprehensive experiments show that TEMPER matches or outperforms existing methods across language modeling and commonsense reasoning tasks, while requiring substantially fewer additional parameters. At eight residual streams, TEMPER achieves the best CORE score while using about $84\%$ fewer additional parameters than mHC, thus showing a stronger performance-parameter efficiency trade-off.

↑