arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11836cs.DBcs.DC

为故障下的复制数据库启用差异化QoS降级

Enabling Differentiated QoS Degradation for Replicated Databases under Failures

Belkis Djeffal, Pierre Bourhis, Romain Rouvoy

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对复制数据库服务,提出在故障下通过PLB中间件实现差异化QoS降级的策略,实验显示其能提升Premium服务的吞吐量保留率、吞吐量并降低延迟。

中文摘要 AI 辅助

弹性通常被视为故障后容量损失的默认应对方式,因为替换副本可补偿故障节点并恢复故障前的服务水平。替换容量既会带来延迟,也需要额外的资源投入,因为必须先配置并同步副本才能处理流量。在预算固定或运行条件受限的情况下,容量恢复不能被视为即时恢复路径。相反,故障处理必须明确在容量减少期间服务如何继续运行。当服务提供差异化服务水平时,容量损失无法统一处理,降级成为服务行为的一部分,需要明确控制减少的容量如何影响每个服务类别,同时不消除预期的差异化。我们研究具有服务类别感知会话的复制数据库服务中的差异化QoS降级,提出一种修复至目标策略,该策略在PLB中实现,PLB是用于服务类别感知路由的PostgreSQL JDBC中间件负载均衡器。当故障停止故障移除部分可用容量时,PLB将健康副本的角色更新为Premium、Mixed和Freemium角色,这保持了副本池的共享状态,同时确保新会话分配继续反映服务类别。我们在两种部署策略(隔离的按服务类别副本池和无优先级感知的共享路由)下,针对单个和级联副本故障对PLB进行评估。结果显示,在Premium侧故障下,PLB将Premium吞吐量保留率的中位数提高了26至28个百分点;在最严重的级联故障阶段,实现了超过2倍的Premium吞吐量;与共享轮询相比,将Premium的p95延迟降低了18.2%。

英文摘要

Elasticity is commonly presented as the default response to capacity loss after failures, since replacement replicas can compensate for failed nodes and restore pre-incident service levels. Replacement capacity entails both delay and additional resource commitment, as replicas must be provisioned and synchronized before they can serve traffic. Under fixed budgets or constrained operating conditions, capacity restoration cannot be treated as the immediate recovery path. Failure handling must instead define how the service continues while capacity remains reduced. When the service exposes differentiated service levels, capacity loss cannot be handled uniformly. Degradation becomes part of the service behavior, requiring explicit control over how reduced capacity affects each class without erasing the intended differentiation. We study differentiated QoS degradation in replicated database services with service-class-aware sessions. We present a repair-to-target policy, implemented in PLB, a PostgreSQL JDBC middleware load balancer for service-class-aware routing. When fail-stop failures remove part of the available capacity, PLB updates the role assignment of healthy replicas into Premium, Mixed, and Freemium roles. This keeps the replica pool shared while ensuring that new session assignments continue to reflect the service class. We evaluate PLB under single and cascading replica failures across two deployment strategies: isolated perclass replica pools and shared, priority-agnostic routing. The results show that PLB improves median Premium goodput retention by 26-28 percentage points under a Premium-side fault, achieves more than 2x higher Premium goodput in the most severe cascading-failure phase, and reduces Premium p95 latency by 18.2% relative to shared round-robin.

补充信息

↑