发表机构
East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对混合DDR--CXL内存数据库,提出五种索引设计并开发自适应数据分布优化,通过分层热度跟踪和收益感知驱逐提升性能,使单一索引设计吞吐量最高提升至1.59倍。
AI 中文摘要
CXL内存扩展了内存数据库的单节点内存容量,但其延迟高于本地DDR,带宽低于本地DDR。混合DDR--CXL数据库需要联合的索引和数据分布设计。不同的索引设计在访问路径和迁移路径上有所差异,并约束了索引和元组的放置,从而在运行时性能、DDR内存效率和系统集成复杂度之间产生权衡。因此,其性能影响难以确定。索引和元组可能具有不同的访问特征:将它们放置在同一层级可能降低DDR效率,而分开管理则增加了跟踪和决策开销。我们提出了基于两种方法的五种索引设计:将两个层级视为统一的数据空间,或分别管理它们。我们分析了它们的权衡及其根本原因,并为数据库显式管理分布的单一索引设计开发了自适应数据分布优化。该方法分别管理索引对象和元组,通过分层热度跟踪和索引对象驻留策略减少元数据开销,并扩展了收益感知准入和随机驱逐,以根据每种类型的预期收益和内存占用确定和调整其放置。我们在混合DDR--CXL数据库原型中实现了这些设计,并使用YCSB进行了评估。在大多数工作负载和系统配置下,采用直接寻址的单一索引设计实现了最佳性能。双索引设计主要在DDR容量受限或访问高度集中时表现更好,并且具有相对较低的集成复杂度。自适应分布优化提高了DDR内存效率和系统性能,使单一索引设计能够实现其非分离管理对应版本吞吐量的高达1.59倍。
英文摘要
CXL memory expands single-node memory capacity for in-memory databases but has higher latency and lower bandwidth than local DDR. Hybrid DDR--CXL databases require joint indexing and data distribution design. Indexing designs differ in access and migration paths and constrain the placement of indexes and tuples, creating trade-offs in runtime performance, DDR memory efficiency, and system integration complexity. Their performance impact is thus difficult to determine. Indexes and tuples may differ in access characteristics: placing them in the same tier may reduce DDR efficiency, while separate management adds tracking and decision overhead. We propose five indexing designs based on two approaches: treating the two tiers as a unified data space or managing them separately. We analyze their trade-offs and underlying causes and develop adaptive data distribution optimization for single-index designs in which the database explicitly manages distribution. The method manages index objects and tuples separately, reduces metadata overhead through hierarchical hotness tracking and an index object residency policy, and extends benefit-aware admission and randomized eviction to determine and adjust their placement according to each type's expected benefits and memory footprints. We implemented the designs in a hybrid DDR--CXL database prototype and evaluated them using YCSB. The single-index design with direct addressing achieves the best performance under most workloads and system configurations. The dual-index design performs better mainly when DDR capacity is limited or accesses are highly concentrated, and has relatively low integration complexity. Adaptive distribution optimization improves DDR memory efficiency and system performance, enabling single-index designs to achieve up to $1.59\times$ the throughput of their counterparts with non-separated management.