AI 中文总结
研究固定任务所需最小Transformer模型容量,通过熵界证明线性注意力替代中令牌混合算子内在秩\(r^*\)是紧密下界,梯度下降可恢复此秩,引入注意力原生内在秩恢复完整熵界结构,绘制数据可预测性边界,将熵界重构为注意力原生容量度量。
AI 中文摘要
神经缩放定律描述了随着模型、数据和计算量的增加损失如何减少,但未回答固定任务所需最小模型容量的问题。我们通过熵界来研究,它是Transformer任务内在容量的谱概念。首先证明在线性注意力替代中,令牌混合算子的内在秩\(r^*\)是紧密下界,梯度下降在标准低秩隐式偏差假设下能恢复此秩。天真地将其转移到真实注意力时失败,通过控制插值阶梯定位原因。接着引入注意力原生内在秩,表明在此定义下线性和softmax注意力的完整熵界结构都得以恢复。最后绘制仅数据可预测性的边界,结果将熵界从事后描述重新构建为具有精确表征可预测前沿的注意力原生容量度量。
英文摘要
Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it? We study this through the Entropic Bound, a spectral notion of task-intrinsic capacity for Transformers. We first prove that, in a linear attention surrogate, the intrinsic rank $r^*$ of the token-mixing operator is a tight lower bound: any rank-deficient model incurs unavoidable excess risk, and the bound is achievable at $r^*$. We further show that gradient descent recovers this rank under standard low-rank implicit-bias assumptions, confirm all three properties empirically, and show $r^*$ is recoverable from data before training. We then ask whether this transfers to real attention. A naive transfer fails, and a controlled interpolation ladder localizes the cause precisely: it is not softmax and not a rank constraint, but the input-conditioned nature of attention's mixing operator, which a static weight kernel cannot summarize. Motivated by this, we introduce an attention-native intrinsic rank -- the minimum query-key kernel rank realizing the task within the attention class -- and show that under this definition the full Entropic Bound structure (deficiency, achievability, recovery) is restored for both linear and softmax attention, with the energy effective rank as the estimator robust to softmax distortion. Finally, we map the boundary of data-only predictability: $r^*$ is exactly recoverable for linear QK attention, even without the value map at scale, while softmax attention admits only partial pre-training recovery due to nonlinear inversion and kernel-value identifiability effects. Our results reframe the Entropic Bound from a post-hoc descriptor into an attention-native capacity measure with a precisely characterized predictability frontier.
Comments11 pages, 1 figure, 5 tables. Code and reproduction scripts included as supplementary material