AI 中文总结
研究针对大语言模型表示坍缩问题,通过频谱分析得出混合效率与信息容量的权衡,提出TRSP方法,采用无参数三角盒机制和长度感知门正则化令牌交互拓扑,实验证明该方法在多方面有显著改进。
AI 中文摘要
大语言模型(LLMs)从根本上受到表示坍缩的限制,这是一个严重降低长上下文性能的瓶颈。现有方法可能陷入同质化坍缩(如注意力汇聚导致秩亏缺)和孤立坍缩(如局部注意力导致上下文断开)这两个极端。通过对注意力动态的频谱分析,得出混合效率(频谱间隙)和信息容量(有效秩)之间的内在权衡。提出拓扑正则化侧路径(TRSP),一种非侵入性架构干预来实现频谱平衡。TRSP采用无参数三角盒机制,由轻量级、长度感知门缩放,以正则化令牌交互拓扑。通过近端耦合保持有效秩和远端传播支持非退化混合,在不改变核心注意力的情况下促进更健康的几何转移算子。实验表明在通用能力和长上下文基准测试中有显著改进。例如在训练长度为8倍的NoLiMa上,TRSP保持83%的准确率,分别比差分变压器和门控注意力高出约30和50个百分点。
英文摘要
Large Language Models (LLMs) are fundamentally limited by representation collapse, a bottleneck that severely degrades long-context performance. We identify that existing approaches risk drifting into one of two pathological extremes: homogenization collapse (e.g., attention sinks causing rank deficiency) and isolation collapse (e.g., local attention causing context disconnection). Through spectral analysis of attention dynamics, we derive an intrinsic trade-off between mixing efficiency (spectral gap) and information capacity (effective rank) that standard mechanisms struggle to balance. To resolve this dilemma, we propose the Topologically Regularized Side-Path (TRSP), a non-invasive architectural intervention that achieves spectral balance. TRSP employs a parameter-free Triangular Box mechanism, scaled by a lightweight, length-aware gate, to regularize the token interaction topology. By integrating proximal coupling to preserve effective rank and distal propagation to support non-degenerate mixing, TRSP promotes a geometrically healthier transition operator without altering core attention. Experiments show significant improvements across general capabilities and long-context benchmarks. Notably, on NoLiMa at $8\times$ the training length, TRSP retains $83\%$ accuracy and surpasses the Differential Transformer and Gated Attention by approximately 30 and 50 percentage points, respectively. Code available at: https://github.com/Eziotao-tyd/TRSP.
Comments22pages, 4 figures, poster of icml 2026