AI 中文总结
Themis是一种基于分析的软件定义硬件预取方案,通过页表项的页级提示过滤无用预取,减少约40%无用请求,可提升BOP等预取器在数据中心工作负载下的性能,无需二进制或ISA变更。
AI 中文摘要
数据缓存未命中是数据中心工作负载中停顿周期的重要组成部分,硬件预取器通过提前获取数据来减少这类停顿,已变得日益复杂。然而,为实现高覆盖率,它们必须激进预取,从而产生大量不准确的访问,浪费内存带宽,这在多租户导致内存带宽成为有限资源的数据中心环境中是个问题。我们观察到,对于数据中心工作负载,不准确的预取可在数据页粒度上有效过滤,且不会牺牲预取覆盖率。但在硬件中存储关于预取有用性的页级元数据成本高昂,因此我们提出一种用于数据预取的新型软硬件接口:软件指导硬件在何处预取,硬件则在感兴趣的区域识别并发出预取请求。我们提出Themis,这是一种实现该新接口的基于分析的硬件预取解决方案。Themis利用页表项中存储的页级提示,在运行时禁用某些数据页的预取器。Themis无需二进制或ISA变更,可在不中断进程执行的情况下优化进程。Themis还与现有预取工作正交,可应用于优化任何硬件预取器。我们的结果显示,Themis能减少约40%的无用预取请求,为所有评估的数据中心工作负载预取器带来加速,包括BOP提升4.1%、SPP+PPF提升3.1%、Pythia提升1.4%。
英文摘要
Data cache misses represent a significant portion of stall cycles in datacenter workloads. Hardware prefetchers that reduce such stalls by fetching data ahead of time have become increasingly sophisticated. However, to achieve high coverage, they have to prefetch aggressively, generating many inaccurate accesses that waste memory bandwidth. This is problematic in datacenter environments where memory bandwidth is a limited resource due to high multi-tenancy. We observe that for datacenter workloads, inaccurate prefetches can be effectively filtered on a data page granularity, without sacrificing prefetch coverage. However, storing per-page metadata about prefetch usefulness in hardware is costly, so we propose a novel hardware-software interface for data prefetching: The software directs the hardware on where to prefetch, and the hardware identifies and issues prefetches in the regions of interest. We propose Themis, a profile-guided hardware prefetching solution that implements this new interface. Themis utilizes page-level hints stored in page-table entries to disable the prefetcher for certain data pages at runtime. Themis requires no binary or ISA changes and can be used to optimize processes without disrupting their execution. Themis is also orthogonal to existing works on prefetching and can be applied to optimize any hardware prefetcher. Our results show that Themis is able to achieve around 40% reduction in useless prefetch requests, resulting in speedup for all the evaluated prefetchers for datacenter workloads, including 4.1% for BOP, 3.1% for SPP+PPF, and 1.4% for Pythia.