arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正则化还是本地化:训练时的KV缓存几何结构在量化下何时起作用

Regularize or Localize: When Training-Time KV-Cache Geometry Pays Under Quantization

Libo Sun, Po-Wei Harn, Zewei Zhang, Peixiong He, Xiao Qin

arXiv 2607.17019首次发表:更新:

发表机构

Auburn University; National Central University(奥本大学; 国立中央大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究LeJEPA的\sigreg在语言模型预训练中能否重塑表示及助于KV缓存量化,通过训练1.1亿参数模型发现,\sigreg能降隐藏状态各向异性,对K和V直接正则化可降缓存各向异性,特定配置下训练干预在量化器比例粗时有用,此为首次训练时分布正则化评估。

AI 中文摘要

我们研究了LeJEPA的反坍缩目标\sigreg在标准自回归语言模型预训练期间是否能重塑表示,以及由此产生的几何结构何时有助于KV缓存量化。我们在10B FineWeb令牌上训练了1.1亿参数的模型,并报告了三个发现。首先,在λ = 0.01时,\sigreg使三个配对种子的隐藏状态成对余弦各向异性降低了38%,困惑度在每对中增加不到0.35%,且无一致的零样本损失。其次,这种变化不会从隐藏状态传播到KV缓存,而在持续训练期间直接对K和V应用\sigreg可使四个检查点的平均缓存各向异性降低94%。最后,在未转换的无对称群量化下,直接的KV正则化是所有三个种子中唯一倾向于逐通道缩放的训练条件,在相同的每通道3位方案下,基线的\dnll是直接正则化模型的4.3至7.9倍。然而,在完整的模拟KIVI风格配置下,所有模型都接近 parity,包括存储开销大致匹配时。在这个1.1亿参数的模型中,当量化器比例较粗时,训练干预有帮助;在测试的令牌局部分组、混合KV缩放和零点的组合下,优势消失。据我们所知,这是首次针对事后缓存量化评估标准KV缓存几何结构的训练时分布正则化。

英文摘要

We study whether \sigreg -- LeJEPA's anti-collapse objective -- can reshape representations during standard autoregressive language-model pretraining, and when the resulting geometry helps \kv-cache quantization. We train 110M-parameter models on 10B FineWeb tokens and report three findings. \textbf{(1)} At $λ{=}0.01$, \sigreg reduces hidden-state pairwise-cosine anisotropy by $38\%$ across three paired seeds. Perplexity increases by less than $0.35\%$ in every pair, with no consistent zero-shot loss. \textbf{(2)} This change does not propagate from hidden states to the \kv cache. Applying \sigreg directly to K and V during continued training, however, reduces mean cache anisotropy by $94\%$ across four checkpoints. A matched continuation without the \kv term leaves cache geometry nearly unchanged, and the frozen-trunk retrofits we tested do not reproduce the effect. \textbf{(3)} Under untransformed symmetric group-free quantization, direct \kv regularization is the only training condition that prefers per-channel scaling in all three seeds, and under that same 3-bit per-channel scheme the baseline incurs $4.3$--$7.9\times$ the directly regularized model's \dnll. Under the full simulated KIVI-style configuration (mixed arrangement, zero-points, grouped scales), however, all models reach near-parity, including when storage overhead is approximately matched. In this 110M regime, the training intervention helps when quantizer scales are coarse; the advantage vanishes under the tested combination of token-local grouping, mixed \kv scaling, and zero-points. To our knowledge this is the first training-time \emph{distributional} regularization of standard \kv-cache geometry evaluated against post-hoc cache quantization.

Comments16 Pages, 4 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑