arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22142cs.DCcs.LG

学习何时优于启发式?Kubernetes调度器评分插件案例研究

When Does Learning Beat Heuristics? A Case Study in Kubernetes Scheduler Score Plugins

发表机构I·拉扎科夫吉尔吉斯国立技术大学
查看机构详情
  • I. Razzakov Kyrgyz State Technical University(I·拉扎科夫吉尔吉斯国立技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Wang Xuying, Zhibek Sarypbekova

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过Kubernetes调度器评分插件案例,对比学习模型与手工启发式方法,发现目标不匹配(逐点回归而非排序目标)是学习模型性能不佳的主因,而非架构问题。

中文摘要 AI 辅助

在生产环境中,Kubernetes调度器中对候选节点进行评分的插件是手工调优的启发式方法。我们探究一个关键问题:基于生产集群轨迹中的真实放置决策训练出的学习评分函数,能否匹配或超越这些启发式方法?如果不能,原因何在?我们为一个广泛使用的调度器模拟器实现了一个基于HTTP的外部评分插件,并评估了两种学习模型(一种基于工程特征构建的随机森林,以及一种基于每个作业的任务依赖图进行编码的图神经网络编码器),这些模型均在大型生产集群轨迹上进行了训练。在标准回归拟合(R^2)指标下,两种模型在四次特征工程迭代中均实现了适度但单调的提升,最终R^2达到约0.042。然而,在调度真正重要的指标——Top-1排名准确率(即模型是否将生产调度器实际选择的机器评为最高分)上,两种学习模型均被一个简单的单特征启发式方法(按空闲CPU排序:74-84%对比任一模型的65-66%)所超越。我们证明,这一差距的最佳解释是目标不匹配:两种模型均使用逐点回归(MSE)而非排名特定目标进行训练,这呼应了学习排序文献中一个长期存在的区分。这与先前关于强化学习训练的调度器在训练目标与部署任务对齐时能优于启发式方法的证据相呼应,表明目标错位(而非架构)是此处的主要障碍。我们进一步报告了一项消融研究,涉及使离线轨迹数据可用所需的占用率重建(朴素特征导致R^2接近0),一项隔离特征丰富度和数据量的受控比较(在两种模型家族之间进行),以及一项关于推理延迟和服务容器内存约束的敏感性分析。代码、数据管道和实验脚本均已发布,以确保可复现性。

英文摘要

Kubernetes scheduler plugins that score candidate nodes are, in production, hand-tuned heuristics. We ask whether a learned scoring function - trained on real placement decisions from a production cluster trace - can match or exceed these heuristics, and if not, why. We implement an external, HTTP-backed scoring plugin for a widely used scheduler simulator and evaluate two learned models (a Random Forest over engineered features, and a graph neural network encoder over per-job task-dependency graphs) trained on a large-scale production cluster trace. Under standard regression fit (R^2), both models improve modestly but monotonically across four feature-engineering iterations, reaching R^2 of about 0.042. However, on the metric that actually matters for scheduling - Top-1 ranking accuracy, whether the model scores the machine the production scheduler actually chose highest - both learned models are outperformed by a trivial single-feature heuristic (rank by free CPU: 74-84% vs. 65-66% for either model). We show this gap is best explained by objective mismatch: both models were trained with pointwise regression (MSE) rather than a ranking-specific objective, echoing a long-standing distinction in the learning-to-rank literature. This parallels prior evidence that RL-trained schedulers can outperform heuristics when the training objective is aligned with the deployment task, suggesting objective misalignment, not architecture, is the primary obstacle here. We further report an ablation of the occupancy reconstruction required to make offline trace data usable (naive features yield R^2 near 0), a controlled comparison isolating feature richness and data volume between the two model families, and a sensitivity analysis of inference latency and serving-container memory constraints. Code, data pipelines, and experiment scripts are released for reproducibility.

↑