arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度神经网络的与宽度无关的可压缩性

Width-Independent Compressibility of Deep Neural Networks

Hong-Yi Wang, Mingze Wang, Liu Ziyin

arXiv 2608.21752首次发表:更新:

发表机构

Princeton University; Peking University; Massachusetts Institute of Technology(普林斯顿大学; 北京大学; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对深度多层感知机证明了与原网络宽度无关的可压缩性定理,提出含导数匹配技术与逐层重加权的构造方法,得出压缩宽度量级公式,为神经网络压缩提供理论支撑。

AI 中文摘要

长期以来,人们已知训练良好的神经网络可被大幅压缩而不影响其性能,这一重要现象仍未被充分理解。我们针对具有解析激活函数的深度多层感知机证明了一个统一可压缩性定理:对于一个深度、宽度固定的教师网络,存在一个同深度的窄网络可近似表示与原网络相同的函数。可达的压缩宽度与原网络宽度显著无关,其量级为$O((\log(1/\varepsilon))^{d_{in}})$,其中$\varepsilon$为误差预算,$d_{in}$为有效输入维度。我们的构造涉及一种感知低维输入的新型导数匹配技术,以及一种保留输入-输出映射的逐层重加权方法。

英文摘要

It has long been known that well-trained neural networks can be compressed very strongly without affecting their performance, an important phenomenon that remains poorly understood. We prove a uniform compressibility theorem for deep multilayer perceptrons with analytic activations. For a deep, wide fixed teacher network, there exists a narrow (same depth) network that approximately represents the same function as the original. The reachable compressed width is strikingly independent of the original width, but is $O((\log(1/\varepsilon))^{d_{in}})$, where $\varepsilon$ is the error budget and $d_{in}$ is the effective input dimension. Our construction involves a novel derivative-matching technique which is aware of the low-dimensional input, and a layer-wise reweighting that preserves the input-output mapping.

Comments11 pages main text, 28 pages total, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑