发表机构
North Carolina State University(北卡罗来纳州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CacheFit是首个将电路级成本纳入搜索目标的LLC替换策略设计循环,应用于MPPPB和SHiP++后,生成的Fit-MPPPB和Fit-SHiP++可大幅降低硬件成本,同时保持或优化性能。
AI 中文摘要
末级缓存(LLC)替换的延迟、功耗与面积成本尚未得到充分研究,这并非因为无法对模拟器进行检测——为功耗模型提供输入的事件计数器是常规操作——而是因为没有任何成本模型以替换策略作为输入。CACTI和McPAT可对缓存或核心等常规结构进行定价,但无法对策略所由构成的小型、形状不规则的阵列及相关逻辑进行定价,且没有任何事件计数能揭示受害者选择路径的逻辑深度。若无此类模型,成本从未成为策略搜索可优化的目标。实践中,策略首先在仿真中因IPC表现被选中,待算法确定后才在寄存器传输级(RTL)降低其成本,因此成本进入设计流程的时间过晚,无法对设计产生影响。本文探究从设计之初就将成本纳入考量会产生何种结果。我们构建了CacheFit,这是首个将电路级成本纳入搜索目标的LLC替换策略设计循环。CacheFit将ChampSim与电路级成本模型HARCOM相连,还与一个可提出策略变体的大语言模型相连。在算法仍处于编写阶段时,就对每个候选方案的延迟、面积、功耗和加速比进行评分。设计者可设定面积、功耗和延迟预算,且每次搜索提出的方案及学到的经验都会以可读文本形式存储,供设计者检查和编辑。我们将CacheFit应用于两个已发表的策略:从功耗和晶体管数量方面成本最高的基准MPPPB,生成了Fit-MPPPB,其存储量减少60.3%,静态功耗降低56.3%,晶体管数量减少58.9%,同时SPEC CPU2017上的IPC提高1.1%,命中率提高7.5个百分点;从IPC最高的基准SHiP++,生成了Fit-SHiP++,其存储量减少56%,静态功耗降低61%,晶体管数量减少56%,IPC降低1.54%,命中率降低2.1个百分点。因此,在设计时就考虑成本的策略,能够以相似或更优的性能实现大幅降低的构建成本。
英文摘要
The latency, power, and area costs of LLC replacement remain poorly explored not because simulators cannot be instrumented--event counters feeding a power model are routine--but because no cost model takes a replacement policy as input. CACTI and McPAT price regular structures, a cache or a core, not the small, oddly shaped arrays and dependent logic a policy is made of, and no event count reveals a victim-selection path's logic depth. Without a model, cost has never been an objective a policy search could optimize. In practice, a policy is first selected in simulation for its IPC, and its cost is reduced later in RTL, after the algorithm is fixed. Cost thus enters the design too late to shape it. This paper asks what happens when cost is present from the start. We build CacheFit, the first policy-design loop for LLC replacement with circuit-level cost inside the search objective. CacheFit connects ChampSim to HARCOM, a circuit-level cost model, and to a large language model that proposes policy variants. Every candidate is scored on latency, area, power, and speedup while its algorithm is still being written. The designer sets the area, power, and latency budgets, and every proposal and every lesson the search learns is stored as readable text that the designer can check and edit. We apply CacheFit to two published policies. From MPPPB, the most expensive baseline in power and transistors, it produces Fit-MPPPB: 60.3% less storage, 56.3% less static power, and 58.9% fewer transistors, with 1.1% higher IPC and a hit rate 7.5 percentage points higher on SPEC CPU2017. From SHiP++, the highest-IPC baseline, it produces Fit-SHiP++: 56% less storage, 61% less static power, and 56% fewer transistors, for 1.54% lower IPC and a hit rate 2.1 points lower. A policy designed with its cost in view can therefore be far cheaper to build at similar or better performance.