发表机构
Argonne National Laboratory; University of Chicago(阿贡国家实验室; 芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究前沿LLM编码智能体能否自动化HPC科学应用的检查点/重启实现,构建含16个MPI应用的基准,采用无人工干预流水线生成41个有效实现,验证其可行性并指出模块化碎片化状态的局限。
AI 中文摘要
高效的检查点/重启支持对于弹性HPC科学应用至关重要,但实现该功能需要大量专业知识:开发者必须识别可恢复状态、选择全局一致的检查点,并在重启过程中保持应用不变量。我们研究前沿大语言模型(LLM)编码智能体是否能自动化这一过程。我们构建了包含16个MPI应用的基准套件,这些应用涵盖不同领域、代码规模和关键状态结构,并采用无人工干预的生成-验证-修订流水线对其进行检查点/重启合成评估。在整个基准测试中,该流水线生成了41个可正常工作的弹性实现。我们的结果表明,当关键状态可见或可通过一致抽象访问时,智能体驱动的弹性工程是可行的:成功运行平均耗时不到1小时,消耗约1500万token,且生成的实现具有可忽略的无故障开销,恢复效率可与人类编写的代码相媲美。然而,模块化和碎片化状态仍是主要限制,部分失败尝试消耗超过1亿token和300分钟仍未生成可正常工作的实现。
英文摘要
Efficient checkpoint/restart support is essential for resilient HPC scientific applications, but implementing it requires substantial expertise: developers must identify recoverable state, choose globally consistent checkpoint points, and preserve application invariants during restart. We study whether frontier LLM coding agents can automate this process. We build a benchmark suite of 16 MPI applications spanning diverse domains, code sizes, and critical-state structures, and evaluate them with a no-human-in-the-loop generate--validate--revise pipeline for checkpoint/restart synthesis. Across the benchmark, the pipeline produces 41 working resilient implementations. Our results show that agent-driven resilience engineering is practical when critical state is visible or accessible through coherent abstractions: successful runs finish in under one hour on average, consume about 15M tokens, and produce implementations with negligible failure-free overhead and recovery efficiency comparable to human-written code. However, modularized and fragmented state remains a major limitation, with some failed attempts consuming over 100M tokens and 300 minutes without producing a working implementation.
Journal refe-Science 2026