arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从试点到生产:跨机构联邦训练与科学人工智能的经验教训

From Pilots to Production: Lessons in Cross-Institutional Federated Training and Artificial Intelligence for Science

Olivera Kotevska, Max Carlson, Yan Gao, Francis Jeanson, Yijiang Li, William Lindskog, Mohammad Naseri, Minseok Ryu, Sahil Tyagi, Jerry Watkins, Feiyi Wang, Ravi Madduri, Kibaek Kim

arXiv 2609.39803首次发表:更新:

发表机构

Oak Ridge National Laboratory; Sandia National Laboratories; Argonne National Laboratory; Flower Labs; Arizona State University; Ontario Brain Institute(橡树岭国家实验室; 桑迪亚国家实验室; 阿贡国家实验室; 花实验室; 亚利桑那州立大学; 安大略脑研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文总结美国能源部等跨机构联邦AI从试点到生产的经验,提出五域就绪框架,强调运维、审计与安全,并指出异步联邦等开放问题,建议成立国际工作组。

AI 中文摘要

许多最有价值的科学数据集无法集中管理:它们具有专有性、受出口管制、属于机密或受数据主权限制。这颠覆了通常的范式:模型必须移动到数据所在之处,使得联邦人工智能(AI)成为开放科学的核心基础设施。我们综合了美国能源部国家实验室、行业部署和开源社区的经验教训,形成了一份关于将联邦AI从试点演示推进到可靠的多站点生产的结构化说明。这些经验教训来自五项并行的工作,涵盖受监管环境中的领导级超级计算机和云基础设施。我们围绕一个改编的“五域就绪框架”和对训练模型之下技术栈的分层视图来组织这些经验教训:数据架构、隐私与安全控制、信任与验证、治理与社会技术因素,以及运维——这些层次决定了试点能否变得可靠。问题从“我们能训练它吗?”转变为“我们能安全地运行它、审计它并更改它吗?”我们报告了在内存效率、可靠性和同步方面的系统经验教训;阐述了隐私、安全和去中心化信任在生产中的要求,以及哪些尚未在生产中得到验证;并指出了跨领域的开放问题(异步联邦、泄漏审计、验证标准和协调的数据契约),我们认为这些问题值得成立一个专门的、国际性的开放科学工作组。

英文摘要

Many of the most valuable scientific datasets cannot be centralized: they are proprietary, export-controlled, classified, or bound by data-sovereignty restrictions. This inverts the usual paradigm: the model must move to the data, making federated artificial intelligence (AI) core infrastructure for open science. We synthesize lessons from U.S. Department of Energy national laboratories, industry deployments, and the open-source community into a structured account of moving federated AI from pilot demonstrations to dependable, multi-site production. The lessons come from five concurrent efforts spanning leadership-class supercomputers and cloud infrastructure in regulated settings. We organize them around an adapted \textit{five-domain readiness frame} and a stratified view of the stack beneath a trained model: data architecture, privacy and security controls, trust and verification, governance and socio-technical factors, and operations, the layers that decide whether a pilot becomes dependable. The question shifts from ``can we train it?'' to ``can we operate it, audit it, and change it safely?'' We report systems lessons in memory efficiency, reliability, and synchronization; set out what privacy, security, and decentralized trust require in production, and what is not yet validated there; and identify cross-cutting open problems (asynchronous federation, leakage auditing, verification standards, and harmonized data contracts) that we argue warrant a dedicated, international, open-science working group.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑