arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个频谱,两种资源:自回归预测中的数据-内存缩放

One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction

Chiwun Yang, Xiaoyu Li

arXiv 2609.13500首次发表:更新:

发表机构

City University of Hong Kong; University of New South Wales(香港城市大学; 新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示自回归预测中数据与内存资源受同一能量频谱支配,提出极小极大定律并验证其在不同因果源及注意力机制中的适用性,实验证实了数据-内存坍缩现象。

AI 中文摘要

需要多少学习到的内存才能从更多数据中获益?我们表明,在正熵自回归检索源中,这两种资源受一个预测能量频谱支配。每个坐标贡献其查询概率乘以其未知logit的平方半径。将所得能量频谱记为$\mu$,我们证明了极小极大定律$\mathfrak R^*_{\rm value}(n,B)\asymp_R \Phi_\mu(n^{-1})+\Phi_\mu(\tau_B), \Phi_\mu(t)=\int\min\{x,t\}\\,\mu(\mathrm dx),$,其中$n$为预测块数,学习状态至多有$2^B$个值。数据设定分辨率$1/n$;内存设定通过最优比特分配达到的水平$\tau_B$。完整曲线还恢复了正频谱。能量-维度配对至关重要:具有相同块能量和块维度边际的两个因果源具有不同的数据和内存指数。一个掩码查询-键注意力头学习路径和值,以显式路由、格式和算术误差实现该定律。进一步的结果给出了指数自适应分配、有限精度实现以及双边算术假设下的计算-精度定律。实验恢复了数据-内存坍缩和耦合指数,解释了路由和分配机制,并检查了六个预训练模型规模上的仅权重量化。

英文摘要

How much learned memory is needed to benefit from more data? We show that the two resources are governed by one predictive-energy spectrum in a positive-entropy autoregressive retrieval source. Each coordinate contributes its query probability times the squared radius of its unknown logit. Writing $μ$ for the resulting energy spectrum, we prove the minimax law $\mathfrak R^*_{\rm value}(n,B)\asymp_R Φ_μ(n^{-1})+Φ_μ(τ_B), Φ_μ(t)=\int\min\{x,t\}\,μ(\mathrm dx),$ for $n$ prediction blocks and a learned state with at most $2^B$ values. Data set the resolution $1/n$; memory sets the level $τ_B$ reached by optimal bit allocation. The complete curve also recovers the positive spectrum. Energy-dimension pairing is essential: two causal sources with identical block-energy and block-dimension marginals have different data and memory exponents. A masked query-key attention head learns the route and values, realizing the law with explicit routing, format, and arithmetic errors. Further results give exponent-adaptive allocation, finite-precision realization, and compute-precision laws under two-sided arithmetic assumptions. Experiments recover the data-memory collapse and coupling exponents, explain the routing and allocation mechanisms, and examine weight-only quantization across six pretrained-model scales.

Comments66 pages, 11 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑