arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于分布式生成式AI的多数据集逆问题求解

Multi-Dataset Inverse Problem Solving with Distributed Generative AI

Daniel Lersch, Steven Goldenberg, Johann Rudi, Markus Diefenthaler, Kevin Brager, Xingfu Wu, Yaohang Li, Nobuo Sato

arXiv 2608.26283首次发表:更新:

发表机构

Thomas Jefferson National Accelerator Facility; Virginia Tech; Argonne National Laboratory; Old Dominion University(托马斯杰斐逊国家加速器装置; 弗吉尼亚理工大学; 阿贡国家实验室; 老自治领大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于SAGIPS框架的分布式生成式AI通用框架,扩展分布式数据并行训练范式至非独立同分布多异构数据集,经多探测器散射实验验证,可鲁棒处理不同数据保真度,适配现实多数据集分析。

AI 中文摘要

从多个异构数据集中提取一组共享的、未知的、无法直接测量的量,是各科学领域普遍面临的挑战。一个典型例子是将来自不同测量、具有不同设置(如 varying detector resolutions,即可变探测器分辨率)的数据集相结合。联合分析此类数据集,而非独立分析或简单合并后分析,对于获得精确且无偏的未知量估计至关重要,但需要谨慎处理数据集异质性,且计算要求高。我们提出了一种通用框架,用于在基于生成式AI的逆问题求解器背景下同时分析多个异构数据集。基于我们近期提出的可扩展异步生成式逆问题求解器(Scalable Asynchronous Generative Inverse Problem Solver,SAGIPS)框架,我们将成熟的分布式数据并行训练范式扩展到非独立同分布数据集,其中每个数据集由同一组未知推理参数控制,但覆盖可用特征空间的不同区域。每个数据集通过自身的前向算子和判别器处理,提供互补约束,共同引导共享生成器实现全局参数一致性。我们利用受多探测器散射实验启发的受控设置验证该方法,提供数值证据表明,我们的框架对不同数据保真度具有鲁棒性,这些数据保真度源于卢瑟福实验中未知的探测器系统误差,且我们展示了其在多GPU领导级计算系统上的扩展行为。结果表明,我们的方法非常适合测量条件随测量变化的现实世界多数据集分析。

英文摘要

Extracting a shared set of unknown, not directly measurable quantities from multiple, heterogeneous datasets is a common challenge across scientific domains. A prominent example is the combination of datasets obtained from different measurements with different settings (e.g. varying detector resolutions). Analyzing such datasets jointly, rather than independently or after naive merging, is essential for obtaining precise and unbiased estimates of the unknowns, but requires careful treatment of dataset heterogeneity and is computationally demanding. We present a generalized framework for simultaneously analyzing multiple heterogeneous datasets in the context of generative AI-based inverse problem solvers. Building on our recent Scalable Asynchronous Generative Inverse Problem Solver (SAGIPS) framework, we extend the well-established distributed data-parallel training paradigm to non-identically distributed datasets, where each dataset is controlled by the same set of unknown inference parameters but covers a different region of the available feature space. Each dataset is processed through its own forward operator and discriminator, providing complementary constraints that collectively guide a shared generator toward global parameter consistency. We validate the approach using a controlled setup inspired by a multi-detector scattering experiment. We provide numerical evidence that our framework is robust to different data fidelities, which arise from unknown detector systematics in the Rutherford experiment, and we show the scaling behavior on multi-GPU leadership computing systems. The results show that our approach is well suited for real-world multi-dataset analyses in which experimental conditions vary across measurements.

Comments23 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑