arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Aneto:利用跨工作负载规律性预测系统性能

Aneto: Predicting System Performance by Exploiting Cross-Workload Regularity

Raul Taranco, Rene Mueller, Michael Giardino

arXiv 2608.07179首次发表:更新:

发表机构

Huawei Technologies Zurich, Switzerland(华为技术苏黎世)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Aneto是一种利用跨工作负载规律性的机械经验回归模型,仅需单次原生运行的硬件计数器即可预测系统性能,其CPI误差比现有最佳一次性预测器低2倍,在内存配置变更预测中表现优异。

AI 中文摘要

预测工作负载对内存技术变更的响应,需要估算每个缓存未命中断处理器的比例。传统上,要准确获取该停滞比例,需进行详细模拟、重复测量或繁重的性能分析。存在一次性替代方案,但会牺牲准确性。我们发现,单次原生运行的硬件计数器足以在无需模拟的情况下推断停滞比例。在涵盖整数、浮点、图和AI基准的100多种不同工作负载中,CPI(每指令周期数)与每指令最大内存停滞的关系在每个微架构上都遵循可预测的模式。Aneto是一种机械经验回归模型,利用这一观察结果。一旦在某台机器上通过一小部分参考工作负载拟合该模型,它就能从单次运行中估算任意新工作负载的性能延迟敏感度,从而支持在任意内存配置下进行一阶CPI预测。在6台机器和2个模拟器上,Aneto的CPI误差比最佳现有一次性预测器低2倍。我们直接针对ARM服务器上的硬件测量验证了预测,从本地DDR到HBM,再到约3倍于基线的内存惩罚,其中中位CPI误差为12.7%,第90百分位为35.9%。在超出直接测量范围的8倍内存延迟外推中,Aneto在Zen 5上与参考模型的一致性在中位为14.6%,第90百分位为41%。此外,Aneto还能为工作负载和架构提供定性见解。

英文摘要

Predicting how a workload responds to a change in memory technology requires estimating how much of each cache miss actually stalls the processor. Obtaining this stall fraction accurately has traditionally demanded detailed simulation, repeated measurements, or heavy profiling. One-shot alternatives exist but sacrifice accuracy. We observe that hardware counters from a single native run suffice to infer the stall fraction without simulation. Across more than 100 diverse workloads spanning integer, floating-point, graph, and AI benchmarks, the relationship between CPI and the maximum memory stall per instruction follows a predictable pattern on each microarchitecture. Aneto is a mechanistic-empirical regression model that exploits this observation. Once fitted on a machine across a small set of reference workloads, the model estimates the performance-latency sensitivity of any new workload from a single run, enabling first-order CPI prediction under any memory configuration. Across six machines and two simulators, Aneto reaches 2x lower CPI error than the best prior one-shot predictor. We validate the predictions directly against hardware measurements on an ARM server, from local DDR to HBM and up to ~3x the baseline memory penalty, where the median CPI error is 12.7% and the 90th percentile 35.9%. At an 8x memory-latency extrapolation beyond the reach of direct measurement, Aneto agrees with a reference model on Zen 5 to within 14.6% at the median and 41% at the 90th percentile. Additionally, Aneto provides qualitative insights into workloads and architectures.

Comments16 pages, 9 figures, 8 tables, accepted to MICRO 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑