arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27540stat.MLcs.ITcs.LGmath.COmath.IT

迈向叠加的数学理论

Towards a mathematical theory of superposition

Michael I. Ivanitskiy, John Jasper, Emily J. King, Dustin G. Mixon

首次发表
浏览论文内容

中文总结 AI 辅助

该研究利用框架理论和压缩感知工具构建神经网络叠加的数学理论,针对超完备字典编码的稀疏二元特征向量,证明了随机支撑和最坏情况支撑设定下的支撑恢复定理,确定了实等角紧框架的精确恢复阈值。

中文摘要 AI 辅助

我们利用框架理论和压缩感知的工具,构建神经网络中叠加的数学理论。在我们的模型中,由活跃特征组成的稀疏二元向量\boldsymbol{x}通过超完备字典\boldsymbol{W}进行编码,特征恢复通过应用带合适偏置向量\boldsymbol{b}的\boldsymbol{\text{ReLU}}(\boldsymbol{W}^\top \boldsymbol{W} \boldsymbol{x} + \boldsymbol{b})来完成。我们针对该模型证明了若干恢复定理:在随机支撑设定下,当期望稀疏度达到\boldsymbol{O}(d/\boldsymbol{\text{log}}\boldsymbol{n})量级时,对于几乎紧的低相干字典,我们建立了高概率支撑恢复的保证;在最坏情况支撑设定下,我们给出了允许支撑恢复的稀疏度水平的精确且可计算的准则。我们将该准则应用于高斯随机矩阵和等角紧框架,对于满足\boldsymbol{n} > \boldsymbol{d} + 1}的实等角紧框架,我们根据相干性确定了精确恢复阈值,该实等角紧框架结果的证明依赖于一种新的刻画——该刻画对框架理论研究者也应具有独立意义——即格拉姆矩阵中符号的分布。

英文摘要

We develop a mathematical theory of superposition in neural networks using tools from frame theory and compressed sensing. In our model, a sparse binary vector \(x\) of active features is encoded through an overcomplete dictionary \(W\), and feature recovery is performed by applying \(\operatorname{ReLU}(W^\top W x+b)\) with an appropriate bias vector \(b\). We prove several recovery theorems for this model. In the random-support setting, we establish high-probability support recovery for nearly tight, low-coherence dictionaries, with guarantees when the expected sparsity is up to order \(d/\log n\). In the worst-case support setting, we give a sharp and computable criterion for which sparsity levels permit support recovery. We apply this criterion to Gaussian random matrices and equiangular tight frames. For real equiangular tight frames with \(n>d+1\), we determine the exact recovery threshold in terms of the coherence. The proof of this result for real equiangular tight frames relies on a novel characterization---which should be of independent interest to frame theorists---of the distribution of signs in the Gram matrix.

↑