arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多机构科学人工智能的隐私基础

Privacy Foundations for Multi-Institutional Scientific Artificial Intelligence

Olivera Kotevska, Sumit Jha, Aurélien Bellet, Rui Hu, Nathaniel D. Bastian, Rafael Ferreira da Silva, Ravi Madduri, Kibaek Kim

arXiv 2609.39787首次发表:更新:

发表机构

Oak Ridge National Laboratory; University of Florida; Inria; University of Nevada; Johns Hopkins University; Argonne National Laboratory(橡树岭国家实验室; 佛罗里达大学; 法国国家数字科学研究所; 内华达大学; 约翰斯·霍普金斯大学; 阿贡国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将科学AI的隐私视为由六个要素定义的保证问题,识别领导级设施特有的元数据泄露和侧信道风险,并给出六个研究优先事项,为跨机构场景提供统一的声明陈述与审计形式。

AI 中文摘要

科学人工智能(AI),涵盖从基础模型(FMs)到联邦数据分析管道,正成为国家实验室、大学、医院和工业合作伙伴之间的共享基础设施。这种合作带来了隐私风险,其自然单位往往是机构的参与、研究策略或技术能力,而非单一记录。差分隐私(DP)、联邦学习(FL)、安全计算、可信执行和溯源各自保护了技术栈的部分环节,但它们的保证很少能在混合信任的机构、访问层级和自主智能体之间组合。本文视角将科学AI的隐私重新定义为由六个要素界定的保证问题:受保护资产、观察者、通道、允许的披露、保证和证据。我们通过一个复合跨机构场景的声明登记表来展示这一框架,并利用它评估模型生命周期。由此产生的两个差距是领导级设施特有的:调度器、分配和遥测元数据暴露了机构的资源态势,而仪器附加的控制回路通过共享加速器上的时序和争用泄露研究策略。我们确定了六个研究优先事项:机构级保证、智能体通信隐私、跨层级信息流、隐私兼容的可复现性、领导级规模核算以及仪器侧信道。其贡献在于提供了一种通用形式,用于陈述、比较和审计那些在科学AI技术栈中原本支离破碎的保证声明。

英文摘要

Scientific artificial intelligence (AI), spanning foundation models (FMs) to federated data-analysis pipelines, is becoming shared infrastructure across national laboratories, universities, hospitals, and industrial partners. This collaboration creates privacy risks whose natural unit is often an institution's participation, research strategy, or technical capability rather than a single record. Differential privacy (DP), federated learning (FL), secure computation, trusted execution, and provenance each protect parts of the stack, but their guarantees rarely compose across mixed-trust institutions, access tiers, and autonomous agents. This perspective recasts privacy for scientific AI as an assurance problem defined by six elements: protected asset, observer, channel, permitted disclosure, guarantee, and evidence. We demonstrate the framing through a claim register for a composite cross-institutional scenario and use it to assess the model lifecycle. Two of the resulting gaps are specific to leadership-class facilities: scheduler, allocation, and telemetry metadata expose an institution's resource posture, and instrument-attached control loops leak research strategy through timing and contention on shared accelerators. We identify six research priorities: institution-level guarantees, agent-communication privacy, cross-tier information flow, privacy-compatible reproducibility, leadership-scale accounting, and instrument side channels. The contribution is a common form for stating, comparing, and auditing claims whose guarantees otherwise remain fragmented across the scientific AI stack.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑