AI 中文总结
研究针对基于RTM缓存因串行访问致移位开销大的问题,提出集间相关性感知的数据放置和替换方案,通过智能分组相关块减少移位开销与缓存能耗,经模拟验证效果显著且扩展性强。
AI 中文摘要
当今以数据为中心的应用程序需要能够随着工作负载增长而扩展,同时保持高性能和能源效率的缓存架构。诸如面积消耗过大和漏电流等基本问题对传统基于SRAM的缓存提出了越来越大的挑战,从而促使人们探索非易失性替代方案。其中,赛道内存(RTM)因其通过纳米线承载顺序磁畴实现的显著存储密度而脱颖而出,这些磁畴可通过畴壁或斯格明子技术进行操纵。尽管有优势,但RTM固有的串行访问会带来相当大的移位开销,导致能耗和延迟增加。本文分析了传统基于RTM的末级缓存中的放置策略,发现当前方法会触发冗余移位操作。为解决此问题,我们引入了一种创新的数据放置和替换方案,智能地对相关块进行分组,确保单次移位不仅能检索目标块,还能使后续块更靠近访问端口对齐。我们使用gem5模拟器和SPEC CPU2017基准测试的模拟结果表明,该方案将移位开销降低了49.0%,缓存能耗降低了35.9%,而对性能的影响可忽略不计。此外,该方案对更长纳米线轨道具有强大的扩展性,可实现更高的缓存密度。
英文摘要
Today's data-centric applications demand cache architectures that can scale with growing workloads while maintaining high performance and energy efficiency. Fundamental issues such as excessive area consumption and leakage power are increasingly challenging traditional SRAM-based caches, thereby motivating the exploration of non-volatile alternatives. Among these, racetrack memory (RTM) stands out due to its remarkable storage density, achieved through nanowires hosting sequential magnetic domains that can be manipulated via domain wall or skyrmion techniques. Despite its advantages, racetrack memory's inherent serialized access introduces considerable shift overhead, leading to elevated energy consumption and latency. In this paper, we analyze the placement strategies in conventional RTM-based last-level caches and identify that current methods trigger redundant shift operations as a shift intended to read one block fails to preposition other blocks subsequently accessed. To resolve this, we introduce an innovative data placement and replacement scheme that intelligently groups correlated blocks, ensuring that a single shift not only retrieves the target block but also aligns subsequent blocks closer to the access port. Our simulation results using the gem5 simulator and the SPEC CPU2017 benchmarks reveal that our scheme reduces shift overhead by 49.0% and cache energy consumption by 35.9% with negligible performance impact. In addition, this scheme exhibits robust scalability to longer nanowire tracks for higher cache density.