arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

计算任务的描述性调度

Descriptive Dispatch of Computational Work

Vanessa Sochat, Daniel Milroy

arXiv 2608.11524首次发表:更新:

AI 中文总结

该研究针对多集群环境下科学工作流的调度挑战,评估了AI/ML调度智能体的可靠性,发现描述性元数据可显著提升作业执行成功率与应用性能。

AI 中文摘要

由AI/ML驱动的智能体正逐渐融入编排流程。任务调度是接收请求、将其转换为工作负载管理器所需格式并成功提交的任务。在多集群环境中运行科学工作流会带来动态作业转换、调度以及向异构集群提交的重大挑战。这些任务非常适合智能体,智能体可接收任务的文本指令、准备作业规格并执行调度。本研究评估了调度智能体在432次运行中的可靠性,测试了四个提示风格下五个特征维度的所有可能组合,该智能体可靠性极高,成功率达97.9%。我们在多集群实验中测试了包含提交、排队、匹配、评分、选择、转换和调度的完整编排流程,发现描述性元数据使220个已提交作业的成功执行率从48%提升至87%,消除了架构不匹配问题,并使10个可测量应用中的5个性能提升了最高3.3倍。

英文摘要

Agents powered by AI/ML are becoming ingrained in orchestration. Dispatch of work is the task of receiving a request, transforming it for a workload manager, and successfully submitting it. Running scientific workflows across multi-cluster environments introduces substantial challenges of dynamic job transformation, dispatch, and submission to heterogeneous clusters. These tasks are well-suited to agents, which can receive textual instructions for work, prepare job specifications, and dispatch. In this work, we assess the reliability of a dispatch agent across 432 runs, testing all possible combinations of five feature dimensions across four prompt styles. The agent is highly reliable (97.9% success). We test a full orchestration to submit, queue, match, score, select, transform, and dispatch in a multi-cluster experiment. We find that descriptive metadata increases successful execution from 48% to 87% of 220 submitted jobs, eliminating architecture mismatch, and improving performance for five of ten measurable applications by up to 3.3x.

Comments9 pages, 5 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑