arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Diamond Agent:联邦HPC资源的智能体化控制即服务

Diamond Agent: Agentic Control of Federated HPC Resources as a Service

Haotian Xie, Junlin Chen, Mingkai Zheng, Yifan Zhu, Minu Mathew, Max Burnette, Yadu Babuji, Volodymyr Kindratenko, Shivaram Venkataraman, Kyle Chard, Ian Foster, Zhao Zhang

arXiv 2609.06181首次发表:更新:

发表机构

Rutgers University; University of Rochester; National Center for Supercomputing Applications; University of Wisconsin–Madison; University of Chicago; Argonne National Laboratory(罗格斯大学; 罗切斯特大学; 国家超级计算应用中心; 威斯康星大学麦迪逊分校; 芝加哥大学; 阿贡国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Diamond Agent通过类型化技能和事件驱动机制,实现跨异构HPC集群的智能工作流执行,将中位额外完成时间降低10.5倍。

AI 中文摘要

高效地聚合和编排跨异构集群的计算能力以支持HPC工作流,面临四个实际挑战:在独立管理的集群间保持工作流上下文、在站点间移动大型数据集、推理站点特定环境和调度器策略,以及利用实时队列和资源状态进行高效任务调度。为此,我们设计了Diamond Agent,一个智能体系统,以类型化技能为接口,实现跨异构集群的HPC工作流智能执行。Diamond Agent提供了一个面向智能体的工作空间和技能,统一了跨站点资源发现、资源规格说明、数据移动、任务执行和结果检索。一个集中的Diamond Agent实例可以操作多个超级计算机,而无需在每个登录节点上单独部署。Diamond Agent将高层智能体动作转换为有效的站点特定执行,通过Globus Transfer移动数据,并利用实时系统能力和队列信息选择可行的放置位置。其事件驱动的延续机制将智能体动作与长时间运行的批处理作业解耦:持久服务监控远程执行,仅在结果或决策相关事件可用时恢复智能体。我们使用了27小时的遥测数据和19轮匹配的多站点提交轮次(包含83个作业,跨越四台生产超级计算机)进行实验。与固定站点基线相比,Diamond Agent将相对于最快观察放置位置的中位额外完成时间从42秒减少到4秒,降低了10.5倍。

英文摘要

Efficiently aggregating and orchestrating computing power across heterogeneous clusters for HPC workflows faces four practical challenges: preserving workflow context across independently administered clusters, moving large datasets between sites, reasoning about site-specific environments and scheduler policies, and exploiting live queue and resource states for efficient task scheduling. To this end, we design Diamond Agent, an agentic system that enables intelligent execution of HPC workflows across heterogeneous clusters with typed skills as the interface. Diamond Agent provides an agent-facing workspace and skills that unify cross-site resource discovery, resource specification, data movement, task execution, and result retrieval. A centralized Diamond Agent instance can operate multiple supercomputers without being deployed separately on each login node. Diamond Agent translates high-level agent actions into valid site-specific executions, moves data through Globus Transfer, and uses live system capability and queue information to select feasible placements. Its event-driven continuation mechanism decouples agent actions from long-running batch jobs: persistent services monitor remote execution and resume the agent only when a result or decision-relevant event is available. We experiment with 27 hours of telemetry and 19 matched multi-site submission rounds comprising 83 jobs across four production supercomputers. Compared with a fixed-site baseline, Diamond Agent reduces the median additional completion time relative to the fastest observed placement from 42 seconds to 4 seconds, a 10.5x reduction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑