arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从重加权到重写:解锁训练数据归因中有影响力样本的干预效应

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

Yuzhang Luo, Chenpeng Wang, Jianhui Chen, Liangming Pan

arXiv 2609.02771首次发表:更新:

发表机构

State Key Laboratory of Multimedia Information Processing, Peking University; YiXin-AILab; Beijing Academy of Artificial Intelligence(北京大学多媒体信息处理国家重点实验室; 翼星人工智能实验室; 北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对训练数据归因中影响力样本的干预价值问题,提出影响引导的响应重写方法,实验发现其效果优于传统重加权,为训练数据归因方法的评估提供了新方向。

AI 中文摘要

训练数据归因(TDA)旨在识别影响模型行为的训练样本,但其干预价值既取决于所选样本,也取决于对样本的修改方式。影响函数(IF)可估计无限小重加权下的行为变化,但在传统基于权重的干预中,IF 所选样本往往比随机选择的样本优势有限。这引发了一个问题:是有影响力的样本缺乏干预价值,还是重加权无法实现其行为潜力。本文引入影响引导的响应重写方法,该方法使用 IF 识别干预目标,并在保持指令固定的情况下,将这些样本的响应替换为与行为一致或行为相反的监督内容。在四个开放权重的大语言模型(LLM)上,我们以认知弃权(epistemic abstention)为主要测试平台,比较了对相同影响力选定样本进行重写和重加权的效果。响应重写产生更强、更持久且双向的行为转变,而对相同样本进行重加权则产生微弱且不一致的效果。进一步分析表明,影响力选定样本比其他选择器提供更大的重写杠杆作用,且变化仍集中在与目标相关的行为上。相同的定性对比也延伸到安全拒绝领域。这些结果区分了影响估计所捕获的局部重加权效应与其所识别样本的更广泛干预杠杆作用,为 TDA 方法的感知干预评估提供了动机。

英文摘要

Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑