发表机构
Algorithmica Solutions(Algorithmica Solutions)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究解耦缓存预取的预测模型与准入策略,发现置信门控的作用超过预测器,在SPEC CPU2017程序上可减少预取、提升准确率且不显著影响DRAM读取量。
AI 中文摘要
学习型缓存预取器通常会与始终发出请求的经典预测器进行评估,这混淆了预测模型与准入策略。我们通过匹配控制变量来解耦这些变量:将相同的准入门应用于含257个参数的在线多层感知机(MLP)和经典步长预测器。神经模型的优势消失了;在随机流量下,MLP与门控步长预测器表现无差异,在大多数常规流上速度更慢。该门本身在架构上独立于预测器具有实用性:在原生ChampSim中的20个SPEC CPU2017程序上,它减少了35%的预取请求,将准确率从11%提升至15%,但DRAM读取量仅变化0.07%,表明代理指标无法预测终端行为。我们证明门关闭时的执行完全复现无预取基线。门比预测器更重要,更好的代理并不意味着更好的终端。
英文摘要
Learned cache prefetchers are typically evaluated against classical predictors that always issue requests, confounding the prediction model with the admission policy. We disentangle these variables with matched controls: the same admission gate is applied to both a 257-parameter online MLP and a classical stride predictor. The neural advantage vanishes; the MLP is indistinguishable from gated stride on random traffic and slower on most regular streams. The gate itself is architecturally useful independent of the predictor: on twenty SPEC CPU2017 programs in native ChampSim, it removes 35% of prefetches and improves accuracy from 11% to 15%, but DRAM reads change by only 0.07% demonstrating that proxy metrics do not predict endpoint behavior. We prove gate-closed execution reproduces the no-prefetch baseline exactly. The gate matters more than the predictor, and better proxies do not imply better endpoints.