发表机构
University of Wisconsin – Madison; University of Texas at Austin; Amazon Web Services; ETH Zurich(威斯康星大学麦迪逊分校; 德克萨斯大学奥斯汀分校; 亚马逊云科技; 苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对内存分层中启发式策略滞后于热集变化的问题,提出MANTA,利用运行时访问特征预测页面有用性并集成轻量级学习模型到ARMS,在模拟CXL和Optane上实现最高5.6倍加速。
AI 中文摘要
内存分层已被用于扩展内存容量,特别是在数据中心中,通过将快速DRAM与较慢的分层(包括CXL附加内存)相结合。其有效性取决于将有用页面保留在快速层中,但现有的启发式策略可能滞后于阶段性或突发性工作负载中变化的热集。为了探索这些局限性,我们引入了ChOMP,一个可扩展的离线优化器,可最小化放置和带宽敏感的迁移成本。然后,我们开发了一个基于跟踪的模拟器,利用此参考来识别在线策略的性能机会。受这些结果的启发,MANTA根据运行时访问特征预测未来页面的有用性,并将轻量级学习模型集成到ARMS中。在模拟CXL的八个工作负载上,MANTA在Linux 6.2和6.18上分别实现了相对于ARMS的几何平均加速比1.12倍和1.08倍(4 GB快速内存);在六个Optane工作负载上,实现了1.69倍。在个别工作负载上,MANTA在Linux 6.2上使用模拟CXL比ARMS快1.25倍,在Linux 6.18上快1.21倍,使用Optane快5.6倍。
英文摘要
Memory tiering has been used to expand memory capacity, particularly in datacenters, by combining fast DRAM with slower tiers, including CXL-attached memory. Its effectiveness depends on keeping useful pages in the fast tier, but existing heuristic policies can lag behind changing hot sets in phased or bursty workloads. To explore these limitations, we introduce ChOMP, a scalable offline optimizer that minimizes placement and bandwidth-sensitive migration costs. We then develop a trace-driven simulator that uses this reference to identify performance opportunities for online policies. Motivated by these results, MANTA predicts future page usefulness from runtime access features and integrates a lightweight learned model into ARMS. Across eight workloads on emulated CXL, MANTA achieves geometric-mean speedups over ARMS of 1.12$\times$ and 1.08$\times$ at 4~GB of fast memory on Linux 6.2 and 6.18, respectively; across six Optane workloads, it achieves 1.69$\times$. On individual workloads, MANTA is up to 1.25$\times$ faster than ARMS with emulated CXL on Linux 6.2, 1.21$\times$ on Linux 6.18, and 5.6$\times$ with Optane.
Comments14 Pages, Accepted to EuroSys 2027