AI 中文总结
针对分离式内存数据库中单面RDMA的局限,提出结合轻量级缓存与高效批处理优化的Lotus,在YCSB基准中吞吐量提升8.2倍、p999延迟降低42.9倍。
AI 中文摘要
RDMA已为分离式内存数据库实现了高速数据访问和低延迟通信。尽管已有多种优化技术被提出以加速该场景下基于RDMA的事务,但由于涉及远程CPU参与,双面RDMA在很大程度上未被充分探索,而单面RDMA更受青睐。然而,单面RDMA的大量使用引入了根本性局限:其有限的API无法表达复杂系统功能,如防止饥饿、基于优先级的调度和抢占,这些都是并发控制协议中的关键功能;此外,索引需要多次网络往返,导致网络放大。在本研究中,我们重新探讨分离式内存数据库场景下单面RDMA与双面RDMA之间的长期争议。我们提出Lotus,它通过利用双面RDMA的丰富功能,结合两项关键优化技术解决了双面RDMA的传统局限——内存服务器中的CPU瓶颈:(1)轻量级缓存;(2)高效批处理。Lotus表明,内存服务器中有限的CPU资源若得到智能利用,可将感知到的弱点转化为显著优势。我们的实验研究显示,在YCSB基准测试中,与最先进的基于单面RDMA的方法相比,Lotus实现了高达8.2倍的吞吐量提升和42.9倍的p999尾部延迟降低。
英文摘要
RDMA has enabled high-speed data access and low-latency communication in disaggregated memory databases. While various optimization techniques have been proposed to accelerate transactions with RDMA in this setting, two-sided RDMA has been largely underexplored in favor of one-sided RDMA due to its remote CPU involvement. However, the heavy use of one-sided RDMA introduces fundamental limitations. Its limited APIs cannot express complex system functions such as starvation prevention, priority-based scheduling, and preemption, which are all critical functions in concurrency control protocols. Moreover, indexing requires multiple network round-trips, causing network amplification. In this work, we revisit the long-standing debate between one-sided RDMA and two-sided RDMA in the context of disaggregated memory databases. We present Lotus, which addresses the conventional limitation of two-sided RDMA, i.e., CPU bottlenecks in memory servers, by leveraging the rich functionality of two-sided RDMA with two key optimization techniques: (1) lightweight caching and (2) efficient batching. Lotus demonstrates that limited CPU resources in memory servers, when intelligently utilized, can transform a perceived weakness into a significant advantage. Our experimental study shows that Lotus achieves up to 8.2$\times$ higher throughput and 42.9$\times$ lower p999 tail latency than state-of-the-art one-sided RDMA-based approaches in YCSB benchmark.