发表机构
Royal Holloway University of London(伦敦大学皇家霍洛威学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述从系统视角整合科学计算工作流中的AI参与,提出连续谱分类,围绕五个系统需求分析代表性应用,并指出动态表示与可复现执行等开放问题。
AI 中文摘要
人工智能(AI)日益嵌入到科学计算工作流中,这些工作流结合了模拟、数据处理、优化、可视化以及实验或观测组件。学习模型可以作为显式的工作流组件,保留持久状态,并且在自适应设置中影响后续计算。现有工作已经对科学工作流管理系统、动态和引导式工作流、AI-HPC耦合模式以及机器学习生命周期进行了描述,尽管这些领域通常被分开讨论。本综述从系统视角将它们整合在一起。我们区分了传统科学工作流、机器学习流水线、AI耦合的高性能计算(HPC)工作流以及更广泛的自动化研究工作流,并提出了一个描述AI参与深度的连续谱,从计算阶段到协同自适应工作流控制。相关的系统需求围绕五个关注点组织:控制与编排;计算与执行;数据与模型状态;可复现性与溯源;以及治理与保证。代表性系统和应用包括AI引导的分子模拟、药物和材料发现、模拟-代理耦合以及分布式自动驾驶实验室。工作流级评估从科学进展、执行成本、数据移动、资源使用、弹性和决策可追溯性方面进行考虑。最后,我们指出了动态工作流表示、状态感知恢复、异构调度、可互操作数据平面、模型介导的决策溯源以及可复现的自适应执行中的开放问题。
英文摘要
Artificial intelligence (AI) is increasingly embedded within scientific computing workflows that combine simulation, data processing, optimisation, visualisation and experimental or observational components. Learned models may serve as explicit workflow components, retain persistent state and, in adaptive settings, influence subsequent computation. Existing work has characterised scientific workflow management systems, dynamic and steered workflows, AI--HPC coupling motifs and the machine-learning lifecycle, although these areas are often discussed separately. This review brings them together from a systems perspective. We distinguish conventional scientific workflows, machine-learning pipelines, AI-coupled high-performance computing (HPC) workflows and broader automated research workflows, and propose a continuum describing the depth of AI participation from a computational stage to co-adaptive workflow control. The associated systems requirements are organised around five concerns: control and orchestration; compute and execution; data and model state; reproducibility and provenance; and governance and assurance. Representative systems and applications include AI-steered molecular simulation, drug and materials discovery, simulation--surrogate coupling and distributed self-driving laboratories. Workflow-level evaluation is considered in terms of scientific progress, execution cost, data movement, resource use, resilience and decision traceability. We conclude by identifying open problems in dynamic workflow representation, state-aware recovery, heterogeneous scheduling, interoperable data planes, model-mediated decision provenance and reproducible adaptive execution.
Comments21 pages, 2 figures, 3 tables