发表机构
Argonne National Laboratory; Oregon State University(阿贡国家实验室; 俄勒冈州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对科学应用HPC执行中LLM智能体面临的凭证安全和一次性生成失败问题,提出ECAS系统,通过边缘控制与验证门控机制实现闭环修复,实验成功率从0/6提升至6/6。
AI 中文摘要
科学应用日益依赖高性能计算(HPC),然而将科学家的高层目标转化为正确的目标规模执行仍然脆弱且劳动密集。大型语言模型(LLM)智能体有望自动化这一过程,但仍存在两个障碍:让云托管模型直接访问HPC会暴露凭证和执行权限,而拒绝访问则要求持续的人工监督;并且当生成的工件在特定站点的HPC环境中失败时,一次性生成无法适应。我们提出ECAS,一种边缘控制的智能体系统,用于在有限人工干预下闭环执行科学计算活动。ECAS将推理、控制和执行分离:云托管的LLM提出计划、工件和修复方案;用户控制的边缘智能体保留凭证、工作流状态和执行权限,同时执行策略和资源约束;HPC系统负责计算。其核心机制是验证门控执行:生成的工件通过静态检查和小规模验证,失败触发基于净化执行反馈的修复,目标规模执行仅在验证和策略检查通过后才被允许。ECAS还利用边缘驻留的专家提炼、站点特定技能库,该库从不向云披露。在初步实验中,针对三个科学应用在两个生产ALCF系统上注入六种故障类型,闭环修复将应用成功率从0/6提高到6/6(对比一次性生成),验证门控阻止了所有三次观察到的目标规模失败,技能条件化将成功率从4/6提高到6/6。这些结果表明,将自适应推理委托给云同时在边缘保留执行控制的可行性。
英文摘要
Scientific applications increasingly rely on high-performance computing (HPC), yet translating a scientist's high-level goal into a correct target-scale execution remains brittle and labor-intensive. Large language model (LLM) agents promise to automate this, but two obstacles remain: granting a cloud-hosted model direct HPC access exposes credentials and execution authority, while withholding it demands continuous human supervision; and one-shot generation cannot adapt when generated artifacts fail in a site-specific HPC environment. We present \textsc{ECAS}, an \textbf{E}dge-\textbf{C}ontrolled \textbf{A}gentic \textbf{S}ystem for closed-loop execution of scientific computing campaigns with limited human intervention. \textsc{ECAS} separates \emph{reasoning}, \emph{control}, and \emph{execution}: a cloud-hosted LLM proposes plans, artifacts, and repairs; a user-controlled edge agent retains credentials, workflow state, and execution authority while enforcing policy and resource constraints; and the HPC system computes. Its core mechanism is \emph{validation-gated execution}: generated artifacts pass static checks and small-scale validation, failures trigger repairs from sanitized execution feedback, and target-scale execution is permitted only after validation and policy checks pass. \textsc{ECAS} also draws on an edge-resident library of expert-distilled, site-specific skills that is never disclosed to the cloud. In preliminary experiments with three scientific applications on two production ALCF systems under six injected fault types, closed-loop repair improves application success from 0/6 to 6/6 over one-shot generation, validation gating prevents all three observed target-scale failures, and skill conditioning improves success from 4/6 to 6/6. These results show the feasibility of delegating adaptive reasoning to the cloud while retaining execution control at the edge.