AI 中文总结
本文针对华为GaussDB数据库进行多项关键改进,使其在30TB TPC-H工作负载下的性能超出已公布最佳结果40%,为提升大规模分析工作负载下的数据库性能提供了有效方案。
AI 中文摘要
GaussDB是华为面向大规模部署及最严苛工作负载设计的顶级数据库系统,属于分布式无共享架构,可处理各类工作负载。本文概述了针对GaussDB的一系列修改,旨在提升其在大规模复杂分析工作负载下的性能。经修改后,其在30TB TPC-H工作负载下的性能超出已公布最佳结果40%。实现该卓越性能的关键改进包括采用流水线执行模型、更快速且可扩展性更强的节点间数据洗牌、利用统一总线与统一远程内存访问;还扩展了基于代价的布隆过滤器放置的支持,并实现了多种布隆过滤器流策略,使其可跨节点使用。
英文摘要
GaussDB is Huawei's premier database system, designed for large-scale deployments and the most demanding workloads. It is a distributed shared-nothing system, capable of handling all types of workloads. This paper outlines a series of modifications to GaussDB aimed at improving its performance on large-scale and complex analytical workloads. After these changes, its performance on the TPC-H workload exceeded the best published result by 40% at 30 TB. The key enhancements to achieve this elite performance include adopting a pipeline execution model, a faster and more scalable inter-node data shuffle, exploiting a unified bus and unified remote memory access. We also expanded the support of cost-based Bloom filter placement and implemented several Bloom filter streaming strategies, enabling their use across nodes.