arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时序数据库的六个基准测试维度

Six Dimensions of Benchmarking Time-Series Databases

Jalal Mostafa, Sandro Melissano, Nicholas Tan Jerome, Suren Chilingaryan, Andreas Kopmann

arXiv 2608.01459首次发表:更新:

AI 中文总结

本文提出时序数据库基准测试框架SciTSv2,从六个维度评估四款不同存储引擎的TSDB,揭示单轴基准掩盖的架构行为,为架构师提供性能诊断与存储引擎选择依据。

AI 中文摘要

时序数据库(TSDB)采用针对特定工作负载特性优化的多样存储架构,导致常规基准测试方法下常不明显的不同性能表现与瓶颈。设计可靠数据后端的系统架构师必须了解哪些存储引擎对其特定流水线高效,哪些未来可扩展性约束风险最低。本文提出SciTSv2基准测试框架,从连接并行度、批量写入、时序规律性、多变量序列、混合工作负载、系统指标六个工作负载维度评估时序数据库。利用SciTSv2系统评估4款代表不同存储引擎的TSDB:InfluxDB(时间结构合并树)、TimescaleDB(基于关系数据库)、ClickHouse(列式)、DataLayerTS(专注常规时序)。结果显示,每个维度都会呈现专用单轴基准测试所掩盖的架构行为,包括依赖规律性的权衡、并发读写间的竞争,以及CPU、I/O、磁盘带宽受限的不同瓶颈。结合细粒度系统指标,SciTSv2为架构师提供将性能结果追溯至底层架构原因的诊断工具,支持基于实证、工作负载特定证据的存储引擎选择。

英文摘要

Time-series databases (TSDBs) employ diverse storage architectures optimized for specific workload characteristics, leading to distinct performance profiles and bottlenecks that are often not apparent under conventional benchmarking approaches. System architects designing robust data backends must understand which storage engines are efficient for their particular pipelines and which exhibit the lowest risk of encountering future scalability constraints. This paper presents SciTSv2, a benchmarking framework that evaluates time-series databases across six workload dimensions: connection parallelism, batch ingestion, time-series regularity, multi-variate series, mixed workloads, and system metrics. We exploit SciTSv2 to systematically evaluate 4 TSDBs representing distinct storage engines: InfluxDB (Time-Structured Merge tree), TimescaleDB (based on relational databases), ClickHouse (columnar), and DataLayerTS (specialized in regular time-series). We show that each dimension surfaces architectural behavior that dedicated, single-axis benchmarks obscure, including regularity-dependent trade-offs, contention between concurrent reads and writes, and distinct CPU, I/O, and disk-bandwidth-bound bottlenecks. Paired with fine-grained system metrics, SciTSv2 gives architects a diagnostic tool for tracing performance outcomes back to their underlying architectural causes, supporting storage engine selections grounded in empirical, workload-specific evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑