arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越静态策略:现代微架构策略间的动态选择

Beyond Static Policies: Dynamic Selection Among Modern Microarchitectural Policies

Yanxin Zhang, Ian McDougall, Junnan Li, Shayne Wadle, Vikas Singh, Karthikeyan Sankaralingam

arXiv 2608.01038首次发表:更新:

AI 中文总结

本文通过组合研究发现微架构策略存在交互效应,提出仅需1比特控制的选择器设计,可高效恢复策略组合的优化空间,为处理器性能提升提供新途径。

AI 中文摘要

现代处理器的性能得益于预取器、预测器、替换规则和调度器等交互策略。这些策略通常被单独评估,但在某一组合中表现最佳的策略,在另一组合中可能表现不佳。为研究这些效应,本文首次对2个L1数据缓存(L1D)预取器、2个L1指令缓存(L1I)预取器和2个L2替换策略,在来自49个SPEC CPU2006和SPEC CPU2017轨迹的490个阶段中开展系统性组合研究。我们按阶段级神谕胜率定义最佳全局静态策略(BGSP),Gaze/Entangling/Mockingjay是BGSP,胜率为33.47%,但平均仍比阶段级神谕低1.33%,其中8个基准程序的52个阶段差距超过2.5%。该优化空间可高度压缩:仅改变L1D预取器的Berti/Gaze组合,其总指令 per 周期(IPC)与8个配置神谕的差距仅为0.039%,将运行时控制降至每20万指令窗口仅1比特。基于这一1比特接口,我们将选择器设计视为信息问题:硬件在选择前能获知什么?我们评估三类选择器:仅使用所选策略IPC的选择器、在任一预取器改变缓存状态前被动监控需求流的选择器、以及能暴露非活动策略胜者信号的理想反事实观察者。主要实用结果是,已执行性能反馈和被动需求监控技术可捕获大部分两策略优化空间,在无需执行或模拟非活动预取器的情况下,恢复62.4%至73.4%的成对神谕差距。反事实研究表明,非活动策略观察必须近乎精确且在一个窗口内可用,才能优于已执行性能或被动需求监控。这些结果提出了一种在微架构策略间自适应的通用方法,作为处理器性能提升的额外途径,区别于结构尺寸调整。

英文摘要

Modern processors gain performance from interacting policies: prefetchers, predictors, replacement rules, and schedulers. These policies are often evaluated one at a time, yet a policy that wins in one stack may lose in another. To study these effects, we present the first systematic composition study of two L1D prefetchers, two L1I prefetchers, and two L2 replacement policies across 490 phases from 49 SPEC CPU2006 and SPEC CPU 2017 traces. We define the best global static policy (BGSP) by phase-level oracle-win frequency. Gaze/Entangling/Mockingjay is the BGSP, winning 33.47% of phases, yet it remains 1.33% below the phase oracle on average, with 52 phases across eight benchmarks losing more than 2.5%. The opportunity is highly compressible: a Berti/Gaze pair that changes only the L1D prefetcher comes within 0.039% aggregate IPC of the eight-configuration oracle, reducing runtime control to one bit per 200K-instruction window. Given that one-bit interface, we frame selector design as an information problem: what can hardware know before choosing? We evaluate selectors that use only chosen-policy IPC, selectors that passively monitor the demand stream before either prefetcher changes cache state, and an ideal counterfactual observer that exposes the inactive-policy winner signal. The main practical result is that both executed-performance feedback and passive demand monitoring techniques capture much of the two-policy opportunity, recovering 62.4% to 73.4% of the pairwise oracle gap without executing or emulating the inactive prefetcher. The counterfactual study shows that inactive-policy observation must be nearly exact and available within one window to improve on executed-performance or passive demand monitoring. These results suggest a general method for adapting among microarchitectural policies as an additional pathway for processor improvement, distinct from structural resizing.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑