发表机构
NVIDIA; University of Wisconsin--Madison(英伟达; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究评估四种ML预测器在微架构结构参数与行为策略场景中的排序表现,发现其在局部反转等关键案例上存在局限,周期级模拟仍不可或缺。
AI 中文摘要
机器学习预测器估计处理器性能的速度远快于周期级模拟。然而,对于设计空间探索而言,有价值的测试不仅是重现常规硬件排序,更在于识别不同硬件配置在单个程序阶段的排名情况。我们在两种设计场景中评估了四种机器学习预测器:\u0026lt;em\u0026gt;结构参数\u0026lt;/em\u0026gt;(SP),即改变发射宽度、重排序缓冲(ROB)大小和缓存容量等硬件资源;以及\u0026lt;em\u0026gt;行为策略\u0026lt;/em\u0026gt;(BP),即改变预取和替换算法。在SP场景中,聚合排序表现良好,但反直觉窗口(CIW)——即预期较慢的配置实际更快的情况——在五组具有明确架构先验的非平局窗口中占比22.4%。这些组间的CIW匹配度仅为23.3%至39.9%,每个点估计值均低于50%的随机严格排序参考值。BP场景呈现出不同的失败情况:真实平局占对窗口的37.8%,大多数严格对的间隔仅为几个周期,且没有任何模型族能可靠击败无特征多数基线。NeuroScalar和SimNet低于该基线,Concorde与基线统计上平局,最佳选定的OneDSE头部仅提升2.1个百分点。准确率主要在大间隔时上升。我们进一步表明,这种失败并非模型容量问题:信息论分析显示,当排序结果取决于指令流中不存在的隐藏微架构状态时,任何基于踪迹的预测器都无法超过仅由可观测输入确定的贝叶斯准确率。因此,高周期或聚合排序准确率可能反映了对简单、高间隔案例的掌握,却遗漏了承载最多架构见解的局部反转,而周期级模拟在此方面仍不可或缺。
英文摘要
Machine-learning predictors estimate processor performance far faster than cycle-level simulation. For design-space exploration, however, the valuable test is not merely reproducing the usual hardware ordering, but identifying how different hardware configurations rank on individual program phases. We evaluate four ML-predictors in two design regimes: \emph{Structural Parameters} (SP), varying hardware resources such as issue width, ROB size, and cache capacity; and \emph{Behavioral Policies} (BP), varying prefetching and replacement algorithms. In the SP regime, aggregate ranking is strong, yet counter-intuitive windows(CIW)---where the configuration expected to be slower is faster---constitute $22.4\%$ of non-tied windows across five pairs with a clear architectural prior. CIW match across these pairs is only $23.3$--$39.9\%$; every point estimate is below the $50\%$ random strict-ordering reference. The BP regime presents a different failure: ground-truth ties cover $37.8\%$ of pair-windows, most strict pairs have margins of only a few cycles, and no model family reliably beats a feature-free majority baseline. NeuroScalar and SimNet fall below that baseline, Concorde is statistically tied with it, and the best selected OneDSE head improves by only $2.1$ percentage points. Accuracy rises mainly at large margins. We further show that this failure is not a matter of model capacity: an information-theoretic analysis reveals that when ranking outcomes depend on hidden microarchitectural state absent from the instruction stream, no trace-based predictor can exceed the Bayes accuracy determined by observable inputs alone. Thus high cycle or aggregate ranking accuracy can reflect mastery of easy, high-margin cases while missing the local reversals that carry the most architectural insight and for which cycle-level simulation remains indispensable.