xTier:面向CXL内存的智能分层
xTier: Intelligent Tiering for CXL-Enabled Memory
浏览论文内容
中文总结 AI 辅助
xTier是一种内核驻留的学习型内存分层系统,利用eBPF和量化MLP在微秒级延迟内评分页面,实现低变动放置,在DRAM:CXL比率1:15及以上时,在18个配置中的14个中达到最快性能,并减少13%-22%的页面移动。
中文摘要 AI 辅助
CXL使能的内存扩展了服务器内存容量,但引入了页面放置问题:操作系统必须决定哪些页面应驻留在DRAM中,哪些应驻留在较慢的CXL内存上。现有系统以两种方式之一进行这种权衡。用户空间控制器支持灵活的策略,但将放置决策暴露于调度器抖动和内核-用户空间交叉开销。内核空间系统避免了这种延迟,但依赖于必须跨工作负载泛化的固定启发式方法。我们提出xTier,一个内核驻留的学习型内存分层系统。xTier将eBPF程序附加到PEBS事件,并使用紧凑的量化MLP在内核内以微秒级延迟对采样页面进行评分。xTier不是对每个候选做出反应,而是收敛到当前工作负载阶段的低变动放置,在收敛后降低采样成本,并在工作负载转变时返回到更高的采样节奏。我们在六个内存密集型工作负载上以DRAM:CXL比率从1:5到1:25评估xTier。优势随着DRAM预算收紧而增长。在1:15及以上,xTier在18个配置中的14个中是最快的系统。在不是最快的情况下,它平均落后最佳基线3.9%。它在实现此性能的同时,几何平均移动页面减少13%,在更紧的比率下减少22%。当工作负载改变阶段时,xTier比任何基线更快、更完全地在DRAM中重建其热集。
英文摘要
CXL-enabled memory expands server memory capacity, but introduces a page-placement problem: the operating system must decide which pages should reside in DRAM and which should reside on slower CXL memory. Existing systems make this tradeoff in one of two ways. Userspace controllers support flexible policies, but expose placement decisions to scheduler jitter and kernel-userspace crossing overhead. Kernel-space systems avoid this latency, but rely on fixed heuristics that must generalize across workloads. We present xTier, a kernel-resident learned memory-tiering system. xTier attaches eBPF programs to PEBS events and uses a compact quantized MLP to score sampled pages inside the kernel at microsecond-scale latency. Rather than reacting to every candidate, xTier converges to a low-churn placement for the current workload phase, reduces sampling cost after convergence, and returns to a higher sampling cadence when the workload shifts. We evaluate xTier on six memory-bound workloads at DRAM:CXL ratios from 1:5 to 1:25. The advantage grows as the DRAM budget tightens. At 1:15 and beyond, xTier is the fastest system in 14 of 18 configurations. Where it is not fastest, it trails the best baseline by 3.9% on average. It reaches this performance while moving 13% fewer pages in geometric mean, and 22% fewer at the tighter ratios. When a workload changes phase, xTier rebuilds its hot set in DRAM faster and more completely than any baseline.
发表机构
- University of Colorado Boulder(科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。