arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10515cs.PFcs.ARcs.DC

PASCAL:面向并行扫描的相位感知共享缓存模型

PASCAL: A Progress Divergence-Aware Shared-Cache Model

  • The Hong Kong University of Science and Technology(香港科技大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhongchun Zhou, Ya Wang, Chengtao Lai, Songtao Mao, Wei Zhang

AI总结:

针对并行扫描模式中多核异步导致的缓存缺失率高的问题,提出相位感知共享缓存模型PASCAL,在执行前预测缺失率,支持大规模设计空间探索,在60种配置上MAPE达13.84%。

AI中文摘要:

在现代AI加速器和GPGPU中,许多并发核心反复访问相同的共享数据。这种模式出现在注意力机制中,不同的查询块共享相同的K/V块;在GEMM中,一行中的每个块读取相同的面板;以及许多其他算子中。我们将这种模式命名为并行扫描。由于该模式中存在大量数据重用,缓存有望捕获尽可能多的数据重用,并大幅减少发送到主存储器的请求,以兼顾性能和能耗。然而,实际上,由于多核固有的异步性,实际的缓存缺失率和DRAM流量可能远高于理想情况。在本文中,我们提出了PASCAL,一种用于并行扫描的共享缓存模型。它感知多核间进度差异的动态特征,将差异与占用率等因素的组合相关联,并在执行前预测缓存缺失率。由于预测不需要目标轨迹、时序或计数器,PASCAL支持在周期精确模拟不切实际的规模上进行设计空间探索,其策略无关的界限表明了任何替换策略都无法避免的流量。在NVIDIA GB10 GPU上,包含不同软件流水线深度、占用率和内存访问数据路径的60种配置数据集中,实现了13.84%的平均绝对百分比误差,而物理波TileSight为44.79%,精确符号SDCM为54.16%。

英文摘要:

In modern AI accelerators and GPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query (Q) tiles share the same key and value (K/V) blocks, GEMM, where every tile in a row reads the same slice, and many other operators. We name this pattern shared cyclic scan. Due to a significant amount of data reuse in this pattern, the cache is expected to capture as much data reuse as possible and largely reduce requests sent to the main memory for both performance and energy consumption concerns. However, in reality, because of the intrinsic asynchrony of multi-cores, the actual cache miss rate and DRAM traffic can be much higher compared to ideal cases. In this paper, we propose PASCAL, a shared-cache model for shared cyclic scans. It calibrates finite-run traffic, which reflects progress divergence, on reference configurations and interpolates it along static program structure such as occupancy to predict the miss rate before execution. PASCAL supports software configuration exploration without target traces or counters at scales where cycle-accurate simulation is impractical, while its analysis gives a sharp $2σ-1$ sufficient capacity condition for preserving LRU sharing. Across held-out scans and GEMM on GB10 and Thor, PASCAL reaches 12.04% balanced fill-equivalent miss-rate MAPE. Replacing TileSight's cache component with PASCAL lowers GB10 GEMM latency MAPE from 18.17% to 13.14%.

↑