基于硬件的准同步共识实现可扩展的数据中心复制
Scalable datacenter replication with mostly-synchronous consensus on hardware
浏览论文内容
中文总结 AI 辅助
该研究提出硬件内可扩展复制方案scarHW,其核心为新型POPUC共识算法,基于FPGA智能网卡实现,可将Redis、Zookeeper的性能提升两个数量级,且故障时零停机,实现高健壮性的可扩展复制。
中文摘要 AI 辅助
分布式进程间的数据一致性复制是涉及著名共识问题的任务,其代价极高且难以扩展,尤其影响具有严格性能要求的数据中心服务。为缓解该问题,我们提出硬件内可扩展复制方案scarHW:一种网卡设计,即使在增加副本数量时也能提升一致性复制的吞吐量和延迟,而当前系统仅能小规模运行或采用宽松的一致性保证。scarHW的核心是我们提出的新型POPUC共识算法,该算法在FPGA智能网卡上实现,以充分利用数据中心可编程网络设备的“准同步”特性。与广泛采用的“准异步”协调协议(如Paxos)或无领导者替代方案不同,POPUC实现了一种称为协作共识的广义共识变体,允许同时做出多个决策,在不损害可用性的情况下实现出色的可扩展性。POPUC在进程崩溃停止以及消息发送/接收遗漏故障(涵盖偶然异步)存在时仍能保持安全性保证,且已在TLA+中进行了形式化规格说明和验证。我们的FPGA原型与现有技术相比,将广泛使用的服务Redis和Zookeeper的吞吐量和延迟提升了两个数量级。基于scarHW的服务在少数副本故障时还能实现零停机,提供了一种高度健壮、线速、可扩展的复制系统。
英文摘要
Consistent replication of data among distributed processes -- a task involving the well-known consensus problem -- is notoriously expensive and hard to scale, affecting especially datacenter services with stringent performance requirements. To mitigate this problem, we introduce scalable replication in-hardware ( scarHW ): a network card design that improves throughput and latency of consistent replication even when increasing the number of replicas, whereas current systems operate at a small scale or with relaxed consistency guarantees. At the heart of scarHW is our novel POPUC consensus algorithm, implemented in an FPGA smartNIC to take full advantage of the "mostly synchronous" behavior of programmable network devices in the datacenter. Unlike widely-adopted "mostly asynchronous" coordination protocols such as Paxos or leaderless alternatives, POPUC implements a generalized variant of consensus dubbed collaborative consensus which allows for several simultaneous decisions, achieving great scalability without compromising availability. POPUC preserves safety guarantees in the presence of process crash-stop and message send/receive omission failures (capturing incidental asynchrony) and has been formally specified and verified in TLA+. Our FPGA prototype improves throughput and latency of widely-used services Redis and Zookeeper by up to two orders of magnitude compared to the state of the art. scarHW-based services also achieve zero downtime upon failure of a minority of replicas, offering a highly-robust, wire-speed, scalable replication system.