AI 中文总结
该研究提出将神经网络压缩为种子与量化潜变量的千字节模型,基于映射网络实现权重再生,在保持准确率的同时大幅减少存储字节,适用于带宽受限的部署场景。
AI 中文摘要
存储和传输训练好的神经网络的成本随其参数数量扩展,这是空中更新、设备端库及其他带宽受限部署的瓶颈。我们研究一种极端模型压缩形式,其中可部署产物不是权重,而是用于再生权重的简短“配方”。基于映射网络(将网络权重表示为紧凑可训练潜变量与固定随机基的非线性函数),我们发现仅需存储潜变量,因为基和初始化中心可从整数种子复现。一个模型成为种子与量化潜变量的结合体,其大小由潜变量维度和位宽决定,而非参数数量。我们将该产物形式化,并引入可扩展至投影无法在内存中容纳的网络的分块种子基。实验中,映射模型的准确率与对每个权重激进量化至几位的相同网络相当,且存储所需字节少得多。达到最激进位宽取决于在循环中对潜变量进行量化微调。结果不依赖于特定随机基,且结构化基可让权重即使对于大型网络也几乎免费再生。
英文摘要
The cost of storing and transmitting a trained neural network scales with its parameter count, a bottleneck for over-the-air updates, on-device libraries, and other bandwidth-bound deployments. We study an extreme form of model compression in which the deployable artifact is not the weights but a short recipe for regenerating them. Building on Mapping Networks, which express a network's weights as a nonlinear function of a compact trainable latent and a fixed random basis, we observe that only the latent need be stored, because the basis and initialization center are reproducible from an integer seed. A model becomes a seed together with a quantized latent, whose size is set by the latent dimension and bit width rather than the parameter count. We formalize this artifact and introduce a seeded block-wise basis that scales to networks whose projection cannot be held in memory. In our experiments, a mapped model is as accurate as the same network quantized aggressively to a few bits per weight, while taking far fewer bytes to store. Reaching the most aggressive bit widths depends on fine-tuning the latent with quantization in the loop. The results do not depend on the particular random basis, and a structured basis lets the weights be regenerated almost for free even for large networks.
Comments13 pages, 6 figures, 15 tables