arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CACHEFORGE:面向性能与硬件效率的LLM引导端到端生成式缓存替换策略

CACHEFORGE: LLM-Guided End-to-End Generative Cache Replacement Policy for Performance and Hardware Efficiency

Kaushal Mhapsekar, Bita Aslrousta, Brijesh Kumar Bhayana, Paula Contreras, Azam Ghanbari, Ethan Goodman, Anna Andriiko, Samira Mirbagher Ajorpaz

arXiv 2610.07668首次发表:更新:

发表机构

North Carolina State University(北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CACHEFORGE通过将大语言模型嵌入硬件感知循环,端到端生成并演化缓存替换策略,在SPEC CPU2006上显著提升命中率与IPC,超越现有基线。

AI 中文摘要

现代缓存替换设计趋于饱和,因为它们运行在固定的表示结构、基于手工特征工程和启发式的预测器,或无法自行生成新决策逻辑的离线模仿模型之上。与此同时,替换行为受到预取、抖动、空间局部性和访问类型行为的因果交互影响,产生了难以手动遍历的巨大设计空间。以往方法通常依赖启发式、参数调优或对离线最优策略的模仿,捕捉相关性而非综合新机制,因此其性能提升往往趋于平台期,并在动态工作负载条件下过拟合。CACHEFORGE是首个通过将大语言模型嵌入受控的硬件感知循环中,端到端演化缓存替换策略的框架。在每次迭代中,LLM提出新的C++替换逻辑,策略在基于轨迹的CRC-2 ChampSim模拟器下进行评估,框架通过奖励塑形、结构检查、动态变异、温度调度和跨策略交叉来强制可行性。这种专为缓存替换策略设计的闭环生成-演化循环,能够发现满足硬件约束的紧凑策略,同时探索超越固定预测器结构的算法变换。在SPEC CPU2006上,CACHEFORGE优于所有CRC-2基线。与MPPPB、ReD、Hawk-eye、SHiP++、LIME和LRU相比,总命中率分别提高27.36%、19.69%、13.72%、13.15%、11.83%和5.73%。在内存密集型工作负载上,与LRU、MPPPB、LIME、ReD、SHiP++和Hawkeye相比,IPC分别提高10.15%、7.89%、6.34%、3.64%、3.12%和2.71%。

英文摘要

Modern cache replacement designs saturate because they operate within fixed representational structures, hand-crafted and heuristic based feature-engineered predictors, or offline imitation models that cannot generate new decision logic on their own. At the same time, replacement is shaped by the causal interaction of prefetching, thrashing, spatial locality, and access-type behavior, producing an enormous design space that is difficult to traverse manually. Prior approaches typically rely on heuristics, parameter tuning, or imitation of an offline optimal policy, capturing correlations rather than synthesizing new mechanisms. As a result, their performance gains often plateau and they overfit under dynamic workload conditions. CACHEFORGE is the first framework to evolve cache-replacement policies end-to-end by embedding a large language model inside a governed hardware-aware loop. In each iteration, the LLM proposes new C++ replacement logic, the policy is evaluated under a trace-based CRC-2 ChampSim simulator, and the framework enforces feasibility through reward shaping, structural checks, dynamic mutation, temperature scheduling, and cross-policy crossover. This closed-loop generation-evolution loop specifically designed for cache replacement policy enables the discovery of compact policies that satisfy hardware constraints while exploring algorithmic transformations beyond fixed predictor structures. Across SPEC CPU2006, CACHEFORGE outperforms all CRC-2 baselines. It improves the total hit rate by 27.36%, 19.69%, 13.72%, 13.15%, 11.83%, and 5.73% over MPPPB, ReD, Hawk-eye, SHiP++, LIME, and LRU, respectively. On memory-intensive workloads, it increases IPC by 10.15%, 7.89%, 6.34%, 3.64%, 3.12%, and 2.71% over LRU, MPPPB, LIME, ReD, SHiP++, and Hawkeye.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑