发表机构
Sichuan University(四川大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有扩散模型归因方法无法区分内部变化的问题,提出CADT方法,通过构建反事实对和动态轨迹描述符实现更精准的训练样本归因,在多类归因任务上优于基线。
AI 中文摘要
扩散模型在图像生成领域已取得显著成功,但要将其输出追溯到单个训练样本仍颇具挑战性。现有归因方法常将特定因素的影响压缩为标量响应,导致无法区分不同的内部变化,这对扩散模型而言尤为受限——扩散模型的语义因素是在去噪过程中通过演化的表示动态性显现的。因此,我们将扩散数据归因重新表述为对因素诱导的内部响应轨迹进行归因。本文提出一种新颖的动态轨迹概念归因方法(CADT),我们认为归因不仅应询问「哪个」样本重要,还应询问其影响在生成过程中是「如何」展开的。具体而言,我们在相同的噪声状态下构建匹配的反事实对,以分离特定因素的表示位移,并将其在去噪过程中的方向和幅度演化建模为动态归因特征。对于每个训练样本和生成查询,CADT会提取分阶段特征向量,并沿去噪过程对其进行整合,形成轨迹描述符。在整个训练集上应用相同的构建方式,可得到一组特定因素的轨迹描述符。随后,利用该描述符的协方差统计量构建协方差感知的半正定核,以校准查询和训练表示,并将校准后的查询轨迹与每个训练轨迹进行比较,从而生成最终的训练样本归因分数。在多个公开数据集上的实验表明,CADT在分层、组合和风格归因任务中,均优于现有的扩散归因基线方法。
英文摘要
Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging. Existing attribution methods often compress factor-specific effects into scalar responses, making distinct internal changes indistinguishable. This is particularly limiting for diffusion models, where semantic factors emerge through evolving representation dynamics during denoising. We therefore reformulate diffusion data attribution as attributing factor-induced internal response trajectories. In this paper, we propose a novel Concept Attribution method through Dynamic Trajectories(CADT). We argue that attribution should therefore ask not only \emph{which} examples matter, but also \emph{how} their influence unfolds during generation. Specifically, we construct matched counterfactual pairs at identical noisy states to isolate factor-specific representation displacements, and model their directional and magnitude evolution across denoising as dynamic attribution signatures. For each training example and generated query, CADT extracts stage-wise feature vectors and integrates them along the denoising process to form a trajectory descriptor. Applying the same construction across the training set yields a bank of factor-specific trajectory descriptors. The covariance statistics of this bank are then used to construct . CADT uses this covariance-aware positive-semidefinite kernel to calibrate the query and training representations, and compares the calibrated query trajectory with each training trajectory to produce the final training-sample attribution scores. Experiments on multiple public datasets show consistent improvements over existing diffusion attribution baselines across hierarchical, compositional, and style attribution.