arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于梯度的数据归因方法中形式重于内容

Form Over Content In Gradient-Based Data Attribution Methods

Sunwoo Kim, Seokwon Jung, Sohyung Kim, Seong Joon Oh, Alice Oh

arXiv 2609.19589首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过独立变化任务与答案格式,证明基于梯度的数据归因方法主要追踪格式相似性而非任务语义,并提出应在两者独立变化的数据上测试此类方法以增强可靠性。

AI 中文摘要

基于梯度相似性的数据归因方法被广泛用于分析和选择大型语言模型的训练数据,但梯度相似性实际衡量的是什么仍存在争议。一些人将其解释为识别与任务相关的技能,而另一些工作则报告表面形式是主要因素。我们通过独立改变任务和答案格式来解决监督微调示例中的这一争议。具体来说,我们以不同的答案格式呈现基准测试,使得数据集可以共享任务而不共享格式,或共享格式而不共享任务。我们发现梯度对齐遵循答案格式,因为共享答案格式的基准对强对齐(去衰减余弦接近0.4),而相同基准以不同答案格式类别呈现时则无对齐(接近0.0)。我们证明这种排序从最早的预训练检查点贯穿到后训练阶段,并跨越模型规模和系列。然后我们分析了LESS(一种用于指令微调的基于梯度的数据选择方法)发布的选样,发现每个目标的选样过度代表了目标自身的答案格式。因此,我们证明基于梯度的归因方法追踪格式相似性多于任务语义,这意味着此类方法以及梯度的语义解释应在答案格式和任务独立变化的数据上进行测试,以获得更强的鲁棒性和可靠性。

英文摘要

Data attribution methods using gradient similarity are widely used to analyze and select training data for large language models, but what gradient similarity actually measures is debated. Some interpret it as identifying task-relevant skills, while other work reports that surface form is the main factor. We resolve this debate for supervised fine-tuning examples by varying task and answer format independently. Specifically, we render benchmarks in different answer formats, such that datasets can share a task without a format or a format without a task. We find that gradient alignment follows the answer format, as benchmark pairs sharing an answer format align strongly (disattenuated cosine near 0.4), while same benchmarks rendered with different answer format classes show no alignment (near 0.0). We demonstrate that this ordering holds from the earliest pretraining checkpoints through post-training, and across model scales and families. We then analyze the released selections of LESS, a gradient-based data selection method for instruction tuning, and find that each target's selections over-represent the target's own answer format. Hence, we demonstrate that gradient-based attribution methods track format similarity more than task semantics, meaning that such methods, as well as the semantic interpretation of the gradient, should be tested on data where answer format and task vary independently for greater robustness and reliability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑