发表机构
Argonne National Laboratory; Fermi National Accelerator Laboratory; RIKEN Center for Computational Science(阿贡国家实验室; 费米国家加速器实验室; 理化学研究所计算科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出在科学软件工程中采用协作式AI工作流,由领域专家制定规范、智能体在确定性编排下编写代码并辅以数值验证和人工审批,实验表明简单作者-审阅者循环在成本上更具优势。
AI 中文摘要
智能体人工智能系统可以在软件工程和科学研究中执行广泛的任务,从编写和翻译代码到运行数据分析和可视化的工作流。在科学计算中,难点在于验证智能体生成的代码既正确又对拥有不同专业领域知识的团队成员来说易于理解。因此,我们认为这些系统最好在协作式团队结构中而非完全自动化的情况下使用。在我们提出的工作流中,领域专家编写规范和计划,智能体在确定性编排模式下运行以编写目标代码。每个阶段结束时都会与参考代码进行数值比较,并要求人工审查和批准后才能进入下一阶段。我们评估了这些工作流在将大型高能物理应用从Fortran翻译为C++时的表现,在不同编排器、设计模式和模型下运行相同任务。在十四项实验中,一个带有强制限制的简单作者-审阅者循环完成的文件数量与多智能体工作流相当,而每个文件的成本约为后者的三分之一。
英文摘要
Agentic artificial intelligence systems can carry out a broad range of tasks in software engineering and scientific research, from writing and translating code to running workflows for data analysis and visualization. In scientific computing, the difficulty is verifying that agent-generated code is both correct and understandable to teams whose members bring different areas of expertise. We therefore argue that these systems are best used within collaborative team structures rather than as full automation. In the workflows we propose, domain experts write the specification and plan, and agents operate under a deterministic orchestration pattern to write the target code. Each stage ends with a numerical comparison against the reference code and requires human review and approval before the next begins. We evaluate these workflows on the translation of a large high-energy physics application from Fortran to C++, running the same task under different orchestrators, design patterns, and models. Across fourteen experiments, a simple author--reviewer loop with enforced limits completed a comparable number of files to a multi-agent workflow at about one-third of the cost per file.