arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31564cs.LG

权重对编码:在神经网络权重中诱导更小的文法

Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights

Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出权重对编码(WeightPE),通过直通估计器内嵌有损Re-Pair压缩器,将文法大小作为显式训练目标,在ViT模型上显著缩小文法规模,仅损失少量准确率。

中文摘要 AI 辅助

我们证明神经网络权重可以被显式微调以接受更小的文法。权重对编码(WeightPE)通过在直通估计器内放置一个有损的Re-Pair压缩器来实现这一点。网络的int8权重被展平为一个字符串,并且在全局L2预算内,近似匹配的Re-Pair模式被精确地设为相等。网络使用重写后的权重进行计算,并通过直通估计器进行训练。与固定大小条目的平坦码本不同,文法提供可变长度的模式,并在更大的模式中层次化地重用它们。在CIFAR-10上微调的ViT-B/16和ViT-L/16的MLP权重上,WeightPE产生的Re-Pair文法大小分别为等效int8 QAT运行产生的文法大小的0.43倍和0.38倍,代价是1.9和1.1个准确率点。这一趋势扩展到不同的文法压缩器(LZ78、SEQUITUR),而网络并未针对这些压缩器进行微调。据我们所知,这是首次将文法大小作为网络权重的显式训练目标。

英文摘要

We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.

发表机构

  • Area Science Park Trieste(的里雅斯特科技园)
  • Sapienza University of Rome(罗马大学)
  • Université Côte d’Azur(蔚蓝海岸大学)
  • Inria(法国国家信息与自动化研究所)
  • Paradigma

机构由 AI 辅助整理,请以论文原文为准。

↑