高性能计算应用程序和工作流的描述性执行
Descriptive Execution of HPC Applications and Workflows
浏览论文内容
中文总结 AI 辅助
研究评估智能框架在高性能计算中的表现,该框架结合大语言模型及资源,能完成扩展研究、作业规范转换和生物科学工作流三项任务,揭示特定任务失败模式,为多集群设置开发提供宝贵信息。
中文摘要 AI 辅助
执行和编排软件组件的方式已从人工编写代码转变为描述性文本。在高性能计算中,这种转变体现在应用编排、工作负载管理以及系统监控和调试等方面。实现任务描述性定义的基础手段是使用具有相关工具功能和资源的大语言模型。结合对这些资源的访问的模型构成了一个自主框架。本文评估了一个智能框架在亚马逊网络服务中通过低延迟网络优化和运行高性能计算扩展研究、在工作负载管理器之间准确转换高性能计算作业规范以及设计和运行整个生物科学工作流的程度。结果表明,该框架完成了所有三项任务,并揭示了特定任务的失败模式。在扩展研究中,智能体部署和优化应用程序,但监控运行作业效率低下;在作业转换中,能高精度转换Slurm和Flux之间的规范;在生物科学工作流中,智能体几乎完全重现了专家编写的变异调用管道。这些信息对于开发由智能体处理调度和转换的多集群设置非常宝贵。
英文摘要
The means to execute and orchestrate software components has changed from human-written code to descriptive prose. In high performance computing, this transition is represented in application orchestration, workload management, and system monitoring and debugging, to name a few. The underlying means to enable descriptive definition of tasks is the use of the Large Language Model with associated tool functions and resources. A combination of a model with access to such resources, modeled in software, encompasses an autonomous framework. As fully automated and agentic frameworks are developed for science, it is important to assess reliability and strategies scoped to specific tasks. In this work, we assess the extent to which an agentic framework can optimize and run an HPC scaling study with a low latency network in Amazon Web Services, accurately transform HPC job specifications between workload managers, and design and run an entire biosciences workflow. We find that the framework completes all three tasks while surfacing task-specific failure modes. In the scaling study, agents deploy and optimize applications but monitor running jobs inefficiently, preferring conservative fixed waits over event subscriptions. In job translation, they convert specifications between Slurm and Flux with high accuracy, with processor-affinity flags the most common error. In the bioscience workflow, the agent reproduces an expert-written variant-calling pipeline almost exactly -- agreeing with the reference call set in 18 of 19 completed runs -- and reaches this result through many distinct yet functionally equivalent workflow implementations. This information is invaluable moving forward to developing multi-cluster setups with scheduling and transformation handled by agents.