arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

蒸馏扩散分数差异用于高效训练数据归因

Distilling Diffusion Score Discrepancy for Efficient Training Data Attribution

Shixuan Liu, Joan Serrà, Kin Wai Cheuk, Jinju Kim, Woosung Choi, Yukara Ikemiya, Wei-Hsiang Liao, Jiaqi W. Ma, Yuki Mitsufuji

arXiv 2609.38776首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; Sony AI; University of Texas at Austin; Sony Group Corporation(伊利诺伊大学厄巴纳-尚佩恩分校; 索尼AI; 德克萨斯大学奥斯汀分校; 索尼集团公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散模型训练数据归因成本高且代理损失不准确的问题,提出基于局部分数差异的TID方法及蒸馏版TIDE,实现毫秒级高效归因,性能优于现有方法。

AI 中文摘要

扩散模型的训练数据归因旨在识别影响生成实例的训练样本,但现有方法要么需要昂贵的逐样本梯度计算,要么需要针对查询的模型优化。此外,大多数方法归因于代理损失的变化,而非实际模型生成行为的变化。我们通过直接使用局部分数差异度量来制定归因,解决了这些局限性,该度量适用于任何扩散变体(包括DDPM、EDM和流匹配),并表明该度量可以在不重新训练的情况下估计,作为预条件梯度相似性。我们将此估计器实例化为基于分数差异的训练数据影响(TID),它使用Kronecker因子曲率来避免随机投影和逐样本梯度存储。然后,我们将TID蒸馏为TIDE,一个仅前向的学生模型,在线训练以从扩散模型的内部激活中复现教师的排名。在CIFAR-10、ArtBench-10和MS-COCO上的反事实评估中,TID匹配或超越了最先进的方法,而TIDE在每次查询成本低四到五个数量级的情况下保留了TID的大部分准确性,在毫秒内归因生成样本,且比生成本身更快。

英文摘要

Training data attribution for diffusion models aims to identify the training samples that influence a generated instance, but existing methods either require costly per-sample gradient computation or query-specific model optimization. Moreover, most methods attribute changes in a proxy loss rather than changes in the actual model's generative behavior. We address these limitations by formulating attribution directly with a local score discrepancy measure, which applies to any diffusion variant (including DDPM, EDM, and flow matching), and by showing that such measure can be estimated without retraining, as a preconditioned gradient similarity. We instantiate this estimator as Training-data Influence via score Discrepancy (TID), which uses Kronecker-factored curvature to avoid random projections and per-sample gradient storage. We then distill TID into TIDE, a forward-only student trained online to reproduce the teacher's rankings from the diffusion model's internal activations. Under counterfactual evaluation on CIFAR-10, ArtBench-10, and MS-COCO, TID matches or outperforms state-of-the-art approaches, while TIDE retains most of TID's accuracy at four to five orders of magnitude lower per-query cost, attributing generated samples in milliseconds and faster than the generation itself.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑