隐形但主导:云OLTP数据库中内核I/O机制的大停顿
Invisible Yet Dominant: Big Stalls of Kernel I/O Mechanisms in Cloud OLTP Databases
浏览论文内容
中文总结 AI 辅助
研究发现云OLTP数据库中内核缓冲I/O因单刷新线程排空不足导致停顿,通过增加设备排空而非带宽,可减少70%节流暂停、降低59%最大延迟并提升23%吞吐量。
中文摘要 AI 辅助
大多数数据库,包括PostgreSQL、RocksDB以及最近的AI KV缓存中间件,依赖缓冲I/O,将写回委托给Linux内核。在云中标准的分布式块存储上,这种委托继承了一个隐藏的瓶颈:每个设备由单个内核刷新线程通过高延迟、浅队列路径进行排空。当排空落后时,脏节流会暂停write()系统调用,甚至必须驱逐脏页的读操作也会停顿。这些停顿对iostat和所有标准计数器不可见。本海报从内核内部观察停顿,使用SteelDB中提出的多卷数据放置作为实验杠杆。eBPF探针在写回和块跟踪点上按发出上下文分离写回,并计数每次节流暂停。在三个配置中,具有相同的预置IOPS和带宽,但分别有1、2和4个设备,我们表明增加排空(而非带宽)可将节流暂停减少70%,将最大事务延迟降低59%,并将吞吐量提高23%。
英文摘要
Most databases, including PostgreSQL, RocksDB, and recent AI KV-cache middleware, rely on buffered I/O, delegating write-back to the Linux kernel. On the distributed block storage standard in the cloud, this delegation inherits a hidden bottleneck: each device is drained by a single kernel flusher thread over a high-latency, shallow-queue path. When the drain falls behind, dirty throttling pauses write() system calls, and even reads that must evict dirty pages stall. These stalls are invisible to iostat and every standard counter. This poster observes the stall from inside the kernel, using the multi-volume data placement proposed in SteelDB as the experimental lever. eBPF probes on writeback and block tracepoints separate write-back by issuing context and count every throttle pause. Across three configurations with identical provisioned IOPS and bandwidth but 1, 2, and 4 devices, we show that adding drains, not bandwidth, cuts throttle pauses by 70%, reduces maximum transaction latency by 59%, and raises throughput by 23%.
发表机构
- NTT, Inc.(日本电信电话株式会社)
机构由 AI 辅助整理,请以论文原文为准。