跨组织孤岛下隐私保护金融欺诈检测的协作合成数据
Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos
查看机构详情
- University of Bayreuth(拜罗伊特大学)
- Fraunhofer FIT(弗劳恩霍夫应用信息技术研究所)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出CollaFuse,一种基于扩散模型的协作合成数据方法,用于跨组织隐私保护的金融欺诈检测,在五个数据集上验证其能更一致地提升下游检测性能,价值源于跨组织结构而非本地保真。
中文摘要 AI 辅助
组织寻求从人工智能中获取分析价值,然而相关数据往往分散在不同组织之间,并受到隐私约束。这一问题在金融欺诈检测中尤为突出,因为罕见的欺诈案例和失衡的本地数据集限制了与决策相关的分析。联邦学习能够在无需直接共享数据的情况下实现协作,但并未解决少数类稀缺的问题。合成数据生成可以提供帮助,然而轻量级方法受限于插值,而生成模型则需要大量数据和计算资源。现有的协作生成方法通常依赖联邦学习,给组织带来了相当大的训练负担。在本文中,我们将CollaFuse作为一种基于扩散模型的协作替代方案用于欺诈检测,并在五个欺诈数据集上对其进行评估。与经典过采样、本地生成基线和集中式扩散基准相比,CollaFuse并未实现最高的本地保真度,但在大多数数据集上更一致地提升了下游欺诈检测性能。这些发现表明,合成数据的分析价值更多来源于可迁移的跨组织结构,而非本地真实性。
英文摘要
Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analytics. Federated learning enables collaboration without direct data sharing but does not resolve minority-class scarcity. Synthetic data generation can help, yet lightweight methods are interpolation-bound, while generative models require substantial data and computation. Existing collaborative generative approaches often rely on federated learning, imposing considerable organization-side training burdens. In this paper, we examine CollaFuse as a collaborative diffusion-based alternative for fraud detection and evaluate it across five fraud datasets. Compared with classical oversampling, local generative baselines, and centralized diffusion benchmarks, CollaFuse does not achieve the highest local fidelity but improves downstream fraud detection more consistently across most datasets. These findings suggest that synthetic data create analytical value less through local realism than through transferable cross-organizational structure.