发表机构
University of North Dakota(北达科他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究彩票票号假说中强彩票票号提取难题,提出双评分方法,通过增强分数空间参数化优化取代逐层稀疏性搜索,相比其他方法显著改进强票号提取,提升性能并降低对稀疏超参数的敏感性。
AI 中文摘要
彩票票号假说提出,大型随机神经网络包含稀疏子网,经过可比训练后可匹配密集模型的性能。更强版本认为,充分超参数化的随机网络包含在任何权重训练之前就已准确的子网。现有理论表明此类强彩票票号存在,但可靠提取仍很困难。我们重新审视了用于提取强票号的冻结权重分数训练方法——边缘弹出,并将逐层稀疏性选择确定为核心瓶颈。我们引入双评分,一种增强的分数空间参数化方法,它通过对扩大的分数张量进行优化来取代逐层稀疏性搜索。我们证明在增强分数空间中的固定密度掩码保留了对所有原始坐标掩码的访问,并且我们表明所得方法可解释为在零增强网络上的边缘弹出。在对照实验中,双评分在强票号提取方面比固定密度边缘弹出和初始化时修剪基线有显著改进,在回绕稀疏训练拓扑的性能上有所提升,并且对稀疏超参数的敏感性明显更低。消融实验表明,增益不仅归因于额外的可训练分数参数,还与诱导有效原始稀疏性的增强分数空间竞争有关。
英文摘要
The lottery ticket hypothesis proposes that large random neural networks contain sparse subnetworks that can match the performance of dense models after comparable training. A stronger version asserts that sufficiently overparameterized random networks contain subnetworks that are already accurate before any weight training. Existing theory establishes that such strong lottery tickets exist, but reliable extraction remains difficult. We revisit edge-popup, a frozen-weight score-training method for extracting strong tickets, and identify layerwise sparsity selection as a central bottleneck. We introduce double-scoring, an augmented score-space parameterization that replaces a layerwise sparsity search with optimization over enlarged score tensors. We prove that fixed-density masking in an augmented score space preserves access to all original-coordinate masks, and we show that the resulting method can be interpreted as edge-popup on a zero-augmented network. In controlled experiments, double-scoring substantially improves strong-ticket extraction over fixed-density edge-popup and pruning-at-initialization baselines, improves on the performance of rewound sparse-training topologies, and exhibits markedly lower sensitivity to sparsity hyperparameters. Ablations show that the gain is not merely due to additional trainable score parameters, but is tied to the augmented score-space competition that induces the effective original sparsity.
Comments46 pages