arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31214cs.AIcs.LG

我们在估计哪种影响?反事实设定在数据归因中的作用

Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出影响估计的分歧源于反事实设定不匹配,而非仅近似误差,通过形式化估计量并实验证明设定选择显著影响归因质量,行为对齐设定可识别目标特定样例。

中文摘要 AI 辅助

估计训练样例对模型行为的影响对于数据调试、数据估值和数据归因至关重要。现有的影响估计器常常产生不相容的排名,这通常被归因于近似误差。我们认为,一个更根本的分歧来源是设定不匹配:影响取决于被归因的行为、对每个训练样例施加的干预,以及将干预映射到模型响应的反事实训练过程。当目标行为需要可处理的替代指标(如查询损失、对数几率或边际)时,这些选择尤为重要。我们将影响形式化为一个反事实估计量,区分不同估计量之间的设定不匹配与估计固定估计量时的近似误差,并根据其隐含的设定来组织代表性估计器。我们进一步推导出一个局部分解,揭示了行为信号、训练信号和反事实参数响应如何相互作用。受控实验表明,不同设定下的精确估计量可以导致不同的排名,而近似误差随着扰动远离其线性化点而增大。在噪声标签检测和LLM归因上的实验表明,设定选择显著影响归因质量,尤其是行为替代指标的选择。行为对齐的设定可以识别出被默认基于损失或基于相似性的设定所掩盖的目标特定训练样例。这些结果确立了设定分析作为解释和比较数据影响估计器的必要第一步。

英文摘要

Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to a model response. These choices are especially important when the target behavior requires a tractable surrogate, such as query loss, a logit, or a margin. We formalize influence as a counterfactual estimand, distinguish specification mismatch across estimands from approximation error in estimating a fixed estimand, and organize representative estimators by their implied specifications. We further derive a local decomposition that exposes how behavior signals, training signals, and counterfactual parameter responses interact. Controlled experiments show that exact estimands under different specifications can induce different rankings, whereas approximation error grows as perturbations move farther from their linearization points. Experiments on noisy label detection and LLM attribution show that specification choices significantly affect attribution quality, especially for the choice of behavior surrogate. Behavior-aligned specifications can identify target-specific training examples obscured by default loss-based or similarity-based specifications. These results establish specification analysis as a necessary first step for interpreting and comparing data influence estimators.

发表机构

  • Singapore Management University(新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑