预测随机低维重参数化何时能训练神经网络
Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks
- Tsinghua University(清华大学)
- The University of Manchester(曼彻斯特大学)
- University of Kentucky(肯塔基大学)
- Miami University(迈阿密大学)
- University of Dayton(代顿大学)
- Institute for Biomedical Informatics, University of Kentucky(肯塔基大学生物医学信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对随机低维重参数化训练神经网络的潜在空间规模问题,提出取向分辨二次主公式,引入RaMaN模型,实现内存优化,实验验证其预测与转换跟踪性能优于取向无关近似。
AI中文摘要:
神经网络通常可通过随机低维重参数化进行训练或微调,其中一个小的潜在向量通过一个冻结的随机映射被映射为完整的参数更新。这引发了一个实际问题:潜在搜索空间需要多大才能达到低损失区域?我们首先将已知的可达性转换表示为等效的锥形式,对于紧致凸目标,其中心为极锥的统计维度。我们的主要理论贡献是一个取向分辨的二次主公式,它能从曲率谱和参考到解的位移轮廓中预测随机切片残差。该公式产生了一个自洽的各向同性取向预测器,在仅保守半径的特例中,可恢复早期的高斯宽度二次界。基于此分析,我们引入了随机映射网络(Random Mapping Networks,RaMaN),它使用结构化哈达玛(Hadamard)或种子再生高斯映射实例化预测的潜在维度。这些构造避免了密集随机映射的O(dP)存储,并将优化器状态内存从O(P)减少到O(d)。我们还开发了无矩阵曲率近似和无扫描维度选择。在受控二次和神经曲率实验中,取向分辨预测器紧密跟踪测量的转换位置,且当位移方向重要时,其性能优于取向无关的近似。端到端实验进一步显示,在图像和语言模型中存在明显的、依赖于协议的训练转换。
英文摘要:
Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. This raises a practical question: how large must the latent search space be to reach a low-loss region? We first express the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone. Our main theoretical contribution is an orientation-resolved quadratic master formula that predicts the random-slice residual from both the curvature spectrum and the reference-to-solution displacement profile. It yields a self-consistent isotropic-orientation predictor and, in a conservative radius-only specialization, recovers the earlier Gaussian-width quadratic bound. Building on this analysis, we introduce Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian maps. These constructions avoid the O(dP) storage of dense random maps and reduce optimizer-state memory from O(P) to O(d). We also develop matrix-free curvature approximations and sweep-free dimension selection. Across controlled quadratic and neural-curvature experiments, the orientation-resolved predictor closely tracks measured transition locations and outperforms orientation-agnostic approximations when displacement direction matters. End-to-end experiments further show sharp, protocol-dependent training transitions across image and language models.