解耦聚合内存优化与索引:一种编译器-运行时方法
Decoupling Disaggregated Memory Optimizations from Indexing: A Compiler-Runtime Approach
- The Chinese University of Hong Kong(香港中文大学)
- Purdue University(普渡大学)
- Microsoft Research(微软研究院)
- Simon Fraser University(西蒙菲莎大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出编译器-运行时框架Nox,可自动将未修改的并发索引转换为适配聚合内存的版本,其生成的索引在真实硬件上性能优异,实现了聚合内存优化的解耦与索引的可移植性。
中文摘要 AI 辅助
聚合内存(DM)将计算和内存解耦为可独立扩展的池,通过较慢的互连而非本地总线连接。这种解耦正是DM具有吸引力的原因——但也意味着每个索引现在必须明确考虑远程内存访问及其相关开销。现有最先进的索引设计通过将远程内存逻辑和优化直接嵌入其核心数据结构和并发控制机制来应对这一问题。因此,为一个索引调优的优化无法被提取并复用至另一个索引,甚至同一索引也无法在不同的DM架构上移植,而无需重新设计。随着硬件和索引需求的发展,这种不断升级的、针对每个索引、每个平台的工程负担是不可持续的。在本文中,我们提出Nox,一种编译器-运行时框架,它通过将未修改的并发索引作为输入,自动生成其聚合内存对应版本,无需修改原始索引逻辑,从而打破这种耦合。编译器层重写索引的LLVM IR,以暴露分配、地址和指针依赖信息,供集中式运行时用于驱动缓存和地址转换。实验表明,Nox生成的B+-树、哈希表和跳表在真实的RDMA和CXL硬件上,在所有测试的工作负载中均能稳健扩展。它们的性能也可优于一些专门的手工索引,并与其他索引相当,尤其是在更接近实际的工作负载上。这些结果表明,当今最快的聚合内存优化无需被锁定在单片手工代码中——编译器-运行时栈可在保留成熟索引实现可扩展性的同时对其进行泛化,且不会因可移植性而牺牲性能。
英文摘要
Disaggregated memory (DM) decouples compute and memory into independently scalable pools, connected over a slower interconnect rather than a local bus. This decoupling is exactly what makes DM attractive--but it also means that every index must now reason explicitly about remote-memory access and its associated optimizations. State-of-the-art index designs respond to this by embedding remote-memory logic and optimizations directly into their core data structures and concurrency control mechanisms. Consequently, an optimization tuned for one index cannot be lifted and reused in another, and even the same index cannot be ported to a different DM architecture without a fresh round of redesign. This escalating, per-index, per-platform engineering burden is unsustainable as hardware and index requirements evolve. In this paper, we present Nox, a compiler-runtime framework that breaks this coupling by taking an unmodified, concurrent index as input and automatically generating its disaggregated-memory counterpart, without touching the original index logic. A compiler layer rewrites the index's LLVM IR to expose allocation, address, and pointer-dependency information that a centralized runtime uses to drive caching and address translation. Empirically, Nox-generated B+-trees, hash tables, and skip lists scale robustly on real RDMA and CXL hardware across every workload tested. They can also outperform some specialized, hand-crafted indexes and match others, especially on workloads that are closer to real-world ones. These results show that today's fastest disaggregated-memory optimizations need not stay locked inside monolithic, hand-crafted code--a compiler-runtime stack can generalize them while preserving the scalability of proven index implementations, without sacrificing it for portability.