CW-Ghost:通过容量窗口进行辅助线程预取的无搜索粒度选择
CW-Ghost: Search-Free Granularity Selection for Helper-Thread Prefetching via Capacity Windows
浏览论文内容
中文总结 AI 辅助
研究辅助线程预取中粒度选择问题,提出CW-Ghost方法,通过离线分析估计缓存行填充量并结合容量预算确定迭代粒度,有界同步限制领先块数,在多平台工作负载上实现加速,相比幽灵线程性能提升,证明缓存容量约束可指导粒度选择。
中文摘要 AI 辅助
辅助线程预取通过在主线程之前执行地址依赖链来隐藏不规则内存访问的延迟。但其有效性取决于辅助线程覆盖的未来迭代范围。固定覆盖范围无法始终适应不同工作负载和处理器,而详尽评估候选配置成本高昂。本文提出CW-Ghost,它通过一次离线分析运行来估计目标区域中每个目标迭代生成的平均需求缓存行填充量。将此估计与缓存容量预算相结合得出容量窗口,确定每个预取块的迭代粒度。此外,有界块级同步限制辅助线程领先主线程运行的块数。在英特尔和AMD CPU平台上评估的14个工作负载实例中,CW-Ghost分别比原始程序实现了1.54倍和1.33倍的几何平均加速。与幽灵线程相比,分别将几何平均性能提高了15.8%和10.8%,并在两个平台的候选集中实现了超过99%的经验最优性能。这些结果表明缓存容量约束可有效指导辅助线程预取粒度的选择。
英文摘要
Helper-thread prefetching hides the latency of irregular memory accesses by executing address dependency chains ahead of the main thread. However, its effectiveness depends on the range of future iterations covered by the helper thread. A fixed coverage range cannot consistently accommodate different workloads and processors, whereas exhaustively evaluating candidate configurations incurs substantial configuration cost. This paper presents CW-Ghost, which uses a single offline profiling run to estimate the average demand cache line fill volume generated per target iteration in a target region. CW-Ghost combines this estimate with a cache capacity budget to derive a Capacity Window, which determines the iteration granularity of each prefetch chunk. In addition, bounded chunk-level synchronization limits the number of chunks by which the helper thread may run ahead of the main thread. Across 14 workload instances evaluated on Intel and AMD CPU platforms, CW-Ghost achieves geometric mean speedups of 1.54x and 1.33x, respectively, over the original programs. Compared with Ghost Threading, it improves geometric mean performance by 15.8% and 10.8%, respectively, while achieving more than 99% of the empirically optimal performance within the candidate set on both platforms. These results demonstrate that cache capacity constraints can effectively guide the selection of granularity for helper-thread prefetching.