arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23087cs.LGcs.AI

神经谱容量:仅从网络规格衡量与设计架构

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

Chenyu Zhu, Ruoyu Zhao, Zhichao Lu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出神经谱容量(NSC),一种仅从架构规格计算的闭式标量,用于衡量架构容量,并配套精确动态规划求解器NSC-DP,在资源约束下全局优化架构,优于现有代理指标,并显著加速剪枝与搜索。

中文摘要 AI 辅助

现代Transformer的设计与压缩都可归结为在预算下分配容量。用于这些决策的标准标量,即参数量(#Params)和浮点运算量(#FLOPs),捕捉了规模与计算量,但未捕捉架构结构:两个具有相同参数预算但不同深度-宽度、头数或前馈网络分配的架构,会获得相同的分数,却表现出不同的行为。我们提出神经谱容量(Neural Spectral Capacity, NSC),这是一个基于每个权重矩阵奇异值谱的闭式标量。在标准随机初始化下,Marchenko-Pastur定律使得NSC仅从架构规格即可计算,无需实例化模型、数据或梯度。其逐层可加结构催生了NSC-DP,一个精确的动态规划求解器,能在CPU上数秒内返回在资源约束下全局最大化NSC的架构——这是现有无训练代理的黑盒搜索无法提供的保证。实验上,在七个Transformer和CNN家族(在FlexiBERT上,对于参数量差异小于10%的配对,NSC的τ=0.505,而#Params降至0.082)的排序中,NSC优于#Params、#FLOPs及代表性的无训练代理;NSC-DP在WikiText-103上发现了一个Transformer-XL架构,在2秒内击败了人工设计的基线;并在无任何校准数据的情况下,将LLaMA-7B剪枝为八个常识推理任务上最佳的5.7B模型,比最强的无训练代理基线快约5900倍。

英文摘要

Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure: two architectures with identical parameter budgets but different depth-width, head, or FFN allocations receive identical scores yet behave differently. We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix. Under standard random initialization, the Marchenko-Pastur law renders NSC computable from the architectural specification alone, with no model instantiation, data, or gradients. Its layer-wise additive structure admits NSC-DP, an exact dynamic-programming solver returning the architecture globally maximizing NSC under resource constraints in seconds on a CPU -- a guarantee that black-box search over existing training-free proxies cannot provide. Empirically, NSC outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families (on FlexiBERT, $τ= 0.505$ on pairs differing in #Params by less than 10%, where #Params collapses to 0.082); NSC-DP discovers a Transformer-XL architecture on WikiText-103 that beats the human-designed baseline in 2 seconds; and prunes LLaMA-7B to the best 5.7B model across eight commonsense reasoning tasks without any calibration data, about 5900x faster than the strongest training-free proxy baseline.

发表机构

  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

↑