网格作业运行时间预测的防泄漏与调度感知机器学习
Leakage-Safe and Scheduler-Aware Machine Learning for Grid Job Runtime Prediction
浏览论文内容
中文总结 AI 辅助
本文提出防泄漏的作业运行时间预测方法,使用提交时属性并采用时间性验证,CatBoost在GWA-T-4数据集上取得最佳效果,基于预测的SJF调度将平均等待时间降低50.92%。
中文摘要 AI 辅助
准确的作业运行时间预测可以改善网格和分布式计算环境中的调度感知资源管理,但预测模型必须在现实的部署约束下进行评估。本文重新审视了GWA-T-4 AuverGrid工作负载轨迹上的CPU突发时间预测,并将其重新定义为防泄漏的预执行作业运行时间预测。我们将目标定义为作业级运行时间,仅使用提交时间属性,排除执行后变量,并在时间性和冷启动设置下评估模型,而不是仅依赖随机交叉验证。我们比较了标准回归器、按时间顺序的历史基线、分类编码策略以及具有原生分类处理的CatBoost。我们进一步增加了时间性超参数调优、运行时间可预测性分析、特征消融、按作业长度进行的错误分析以及一个最小调度模拟。经过时间性验证调优后,CatBoost在保留的时间性测试集上取得了最强的部署导向结果,R^2=0.239,MAE=27,019,RMSE=46,587,LogMAE=2.646。对全部69,523个保留的时间性测试作业进行的单服务器模拟表明,基于预测的SJF相对于FCFS将平均等待时间减少了50.92%。结果表明,随机分割评估高估了性能,分类原生提升改善了时间性泛化,而长作业低估仍然是一个与调度器相关的挑战。
英文摘要
Accurate job runtime prediction can improve scheduling-aware resource management in grid and distributed computing environments, but prediction models must be evaluated under realistic deployment constraints. This paper revisits CPU burst time prediction on the GWA-T-4 AuverGrid workload trace and reformulates it as leakage-safe pre-execution job runtime prediction. We define the target as job-level runtime, use only submission-time attributes, exclude post-execution variables, and evaluate models under temporal and cold-start settings rather than relying only on random cross-validation. We compare standard regressors, chronological historical baselines, categorical encoding strategies, and CatBoost with native categorical handling. We further add temporal hyperparameter tuning, runtime predictability analysis, feature ablation, error analysis by job length, and a minimal scheduling simulation. After temporal-validation tuning, CatBoost achieves the strongest deployment-oriented result with R^2=0.239, MAE=27,019, RMSE=46,587, and LogMAE=2.646 on the held-out temporal test set. A single-server simulation over all 69,523 held-out temporal test jobs shows that prediction-informed SJF reduces average waiting time by 50.92% relative to FCFS. The results show that random-split evaluation overestimates performance, categorical-native boosting improves temporal generalization, and long-job underestimation remains a scheduler-relevant challenge.
发表机构
- Augustana College(奥古斯塔纳学院)
- Florida International University(佛罗里达国际大学)
- Dakota State University(达科他州立大学)
机构由 AI 辅助整理,请以论文原文为准。