arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FrED:通过领域知识图谱基础进行外部数据影响估计

FrED: External Data Influence Estimation via Domain Knowledge Graph Grounding

Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose

arXiv 2607.21615首次发表:更新:

发表机构

National Centre for Scientific Research “Demokritos”; University of Glasgow(国家科学研究中心“德谟克利特”; 格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对生成式AI训练数据归因问题,提出在黑盒设置下运行的概率框架,融合连续特征相似度与领域特定知识图谱,在艺术图像合成和天气预报等领域评估,证明其有效性,为外部数据影响分析提供高效可解释机制。

AI 中文摘要

生成式人工智能的快速部署凸显了训练数据归因的迫切需求,以确保透明度和问责制。当前参数方法需访问模型权重,计算成本高,基于相似度的方法忽略深层结构上下文。我们提出一种全新概率框架,在黑盒设置下运行,融合连续特征相似度与离散领域特定知识图谱,确保归因基于结构现实。在抽象艺术图像合成和高维物理天气预报两个领域评估该框架,广泛基准测试证明其有效性。在艺术领域表现出色,在环境预测中还进行了跨域可行性案例研究。该方法无需访问内部模型,为事后影响分析和领域基础检索提供了高效、可解释机制。

英文摘要

The rapid deployment of generative AI has amplified the critical need for Training Data Attribution to ensure transparency and accountability. However, current parametric approaches require computationally prohibitive access to model weights, while similarity-based methods ignore deep structural context. We propose a novel probabilistic framework that operates entirely in a black-box setting. Our method fuses continuous feature similarities with discrete, domain-specific Knowledge Graphs (KGs). This approach ensures the attribution is grounded in structural reality, explicitly rewarding highly specific historical samples while preventing generic background data from dominating the results. We evaluate our framework across two distinct domains where linking outputs to data and domain context is inherently complex: abstract artistic image synthesis and high-dimensional physical weather forecasting. Extensive benchmarking demonstrates the robust efficacy of our approach. In the artistic domain, it achieves a strong Linear Datamodeling Score that exceeds standard black-box similarity baselines, while closing much of the gap to gradient-based estimators. We additionally present a cross-domain feasibility case study in environmental forecasting, where we use domain KGs to retrieve physically consistent historical analogs for regional flood forecasts, improving geographic localisation over a latent-only baseline. Operating entirely without internal model access, our approach provides an efficient, interpretable mechanism for post-hoc influence analysis and domain-grounded retrieval.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑