arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30081cs.LGcs.CV

在流匹配模型中将生成样本溯源至训练数据簇

Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models

  • Forschungszentrum Jülich(于利希研究中心)
  • Technical University Dortmund(多特蒙德工业大学)
  • Reichman University(赖希曼大学)
  • Lamarr Institute(拉马尔研究所)
  • Institute for AI in Medicine, University Hospital Essen(埃森大学医院医学人工智能研究所)
  • Helmholtz AI(亥姆霍兹人工智能)
  • University of Cologne(科隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Rania Briq, Ohad Fried, Michael Kamp, Stefan Kesselheim

AI总结:

本研究针对流匹配模型,提出混合分析-学习方法推导簇级轨迹归因分数,经实验验证其在部分指标具竞争力,揭示流匹配归因受语义相似性等多因素影响。

AI中文摘要:

理解哪些训练样本会影响生成图像,是生成建模中的重要问题。在流匹配模型中,训练样本通过生成轨迹上的速度场影响生成图像;移除样本以检验其反事实影响会改变速度场,而速度场的局部变化对最终图像的影响,取决于该变化在轨迹中的传播情况,因此速度场的局部变化不一定能预测最终的反事实效果。本研究采用分析与学习相结合的混合方法,研究流匹配模型中的归因问题,以此推导簇级别的基于轨迹的归因分数。我们使用独立重新训练的留一数据簇(LOO)模型评估这些归因分数,并在两种不同的流匹配潜在空间中与多个归因基线进行对比实验。实验结果表明,语义相似性是一种强基线;而闭式形式的基于轨迹的归因在部分指标上具有竞争力,且无需反事实重新训练或模型梯度。我们的结果显示,流匹配中的归因不仅取决于与训练样本的语义相似性,还取决于潜在表示、轨迹动态以及影响向最终输出的传播方式。

英文摘要:

Understanding which training samples influence a generated image is an important problem in generative modeling. In flow matching, training samples influence the generated image through the velocity field along the generation trajectory. Removing samples to examine their counterfactual influence changes the velocity field, and the resulting effect on the final image depends on how the change propagates through the trajectory. Consequently, local changes in the velocity field do not necessarily predict the final counterfactual effect. This work investigates attribution in flow-matching models through a hybrid analytical--learned approach, and uses it to derive trajectory-based attribution scores at the cluster level. We evaluate these attribution scores using independently retrained leave-one-cluster-out (LOCO) models, and compare with several attribution baselines using two different flow-matching latent spaces. Our experiments show that semantic similarity constitutes a strong baseline, while the closed-form trajectory-based attribution is competitive in some metrics without requiring counterfactual retraining or model gradients. Our results show that attribution in flow matching depends not only on semantic similarity to training samples, but also on the latent representation, trajectory dynamics, and how influence is propagated to the final output.

↑