AI 中文总结
该研究提出C2P-Cache,通过将GPU L1远程命中搜索转化为轻量级过滤-确认流程,减少冗余L2访问,使GPU IPC最高提升49.7%,平均提升23.5%。
AI 中文摘要
现代GPU依赖每个流式多处理器(SM)私有的L1缓存和共享的L2缓存,但这种架构会隐藏跨SM的缓存复用:即使请求的缓存行已存在于对等L1缓存中,L1缺失(miss)通常仍会被转发到L2,导致冗余的L2访问。此前的GPU L1共享设计尝试通过精确或宽泛的远程命中搜索来恢复这种复用,但随着更多缓存参与且更多缺失并发到达,这类设计的可扩展性越来越差,还会在高并发下干扰关键的L1缺失路径。我们发现,消除冗余L2访问并不需要芯片范围内私有L1内容的精确知识,仅需要足够的可见性来精准缩小候选缓存的范围,再将精确确认留给数量少得多的L1缓存即可。基于这一见解,我们提出C2P-Cache,这是一种可扩展的GPU L1共享机制,将远程命中发现从芯片范围的精确搜索问题转化为轻量级的过滤-确认流程。C2P-Cache维护基于紧凑布隆过滤器(Bloom filter)的私有L1标签快照,执行并行芯片范围的候选过滤,并仅选择性探测少量可能的对等缓存。为维持高并发,C2P-Cache将过滤组织为对分块且复制的快照矩阵的位切片匹配,能高效并行处理大量并发缺失,且不会干扰正常的L1访问。在广泛的GPU工作负载中,C2P-Cache的每周期指令数(IPC)最高提升49.7%;对于高远程L1复用且对L2延迟敏感的应用,平均提升23.5%,证明轻量级可扩展过滤能以适度开销有效释放跨SM复用的潜力。
英文摘要
Modern GPUs rely on private per-SM L1 caches and a shared L2 cache, but this organization obscures cross-SM reuse: an L1 miss is typically forwarded to L2 even when the requested line already resides in a peer L1 cache, leading to redundant L2 access. Prior GPU L1-sharing designs attempt to recover such reuse through exact or broad remote-hit searches, which become increasingly difficult to scale and can interfere with the critical L1 miss path under high concurrency. %miss handling as more caches participate and more misses arrive concurrently. We observe that eliminating redundant L2 accesses does not require exact, chip-wide knowledge of private L1 contents. Instead, it requires only sufficient visibility to sharply narrow down a small set of candidate caches, leaving exact confirmation to a much smaller number of L1s. Based on this insight, we propose C2P-Cache, a scalable GPU L1-sharing mechanism that transforms remote-hit discovery from a chip-wide exact search problem into a lightweight filtering-and-confirmation process. C2P-Cache maintains compact Bloom-filter-based snapshots of private L1 tags, performs parallel chip-wide candidate filtering, and selectively probes only a small number of likely peer caches. To sustain high concurrency, C2P-Cache organizes filtering as bit-sliced matching over a banked and replicated snapshot matrix, enabling efficient, parallel processing of many concurrent misses without interfering with normal L1 accesses. Across a wide range of GPU workloads, C2P-Cache improves instructions per cycle (IPC) by up to 49.7\% and by 23.5\% on average for applications with high remote-L1 reuse and strong sensitivity to L2 latency, demonstrating that lightweight, scalable filtering can effectively unlock cross-SM reuse with modest overhead.
CommentsAccepted at the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)