arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37842cs.LGcs.AI

缩放影响函数在LLMs中通过特征基校正的一位梯度投影

Scaling Influence Functions in LLMs through Eigenbasis-Corrected One-Bit Gradient Projection

Jaeseung Heo, J Rosser, Dongwoo Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出EOGP方法,通过EK-FAC降维、PCA学习压缩方向及一位量化,在固定存储预算下压缩训练梯度以高效计算影响函数,在GPT-2和OLMo 2上以更少存储达到或超越基线精度。

中文摘要 AI 辅助

影响函数估计单个训练示例如何影响大型语言模型(LLMs)的行为。分析训练数据如何影响LLM的不同行为涉及重复的影响计算。重用存储的训练梯度降低了计算成本,但在LLM规模下存储完整梯度是极其昂贵的。我们研究如何压缩这些梯度,同时保留对存储时未知的未来查询的影响估计。通过最坏情况分析,我们刻画了最优固定维度线性表示,并提出了特征基校正的一位梯度投影(EOGP)以在规模上近似它。具体来说,EOGP使用EK-FAC降低梯度维度,然后在保留的子空间内应用PCA从训练梯度中学习压缩方向。然后我们对得到的坐标应用一位量化,允许在固定存储预算内保留更多坐标。在GPT-2上,EOGP预测重训练结果比评估的压缩基线更准确,同时使用其每个示例存储的十六分之一。在从1B到32B参数的OLMo 2 SFT模型上,EOGP与分配每个示例超过100倍存储的基线保持竞争力。

英文摘要

Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients reduces the computational cost, but storing full gradients is prohibitively expensive at LLM scale. We study how to compress these gradients while preserving influence estimates for future queries that are unknown at storage time. Through a worst-case analysis, we characterize the optimal fixed-dimensional linear representation and propose eigenbasis-corrected one-bit gradient projection (EOGP) to approximate it at scale. Specifically, EOGP uses EK-FAC to reduce gradient dimensionality, then applies PCA within the retained subspace to learn compression directions from the training gradients. We then apply one-bit quantization to the resulting coordinates, allowing more coordinates to be retained within a fixed storage budget. On GPT-2, EOGP predicts retraining outcomes more accurately than the evaluated compression baselines while using one-sixteenth of their per-example storage. On OLMo 2 SFT models from 1B to 32B parameters, EOGP remains competitive with the baselines allocated over 100 times as much storage per example.

发表机构

  • POSTECH(浦项科技大学)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

↑