arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

简单读取的背后:理解 Linux I/O 栈中的工作与等待

Behind a Simple Read: Understanding Work and Waiting in the Linux I/O Stack

Yang Shen, Kai Lu, Min Xie, Huijun Wu, Zhenwei Wu, Wenzhe Zhang

arXiv 2610.10137首次发表:更新:

发表机构

National University of Defense Technology(中国人民解放军国防科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过分解 Linux 缓冲读取路径,提出跨层等待链模型,量化设备等待并验证优化机会,为分层 I/O 栈性能决策提供可测量基础。

AI 中文摘要

简单的读取接口提供了统一的功能语义,但其性能行为并非同样简单或可预测。我们以 Linux 中的同步大缓冲区读取作为观察窗口,端到端地分解缓冲读取路径,区分工作量、处理时间和关键路径暴露。我们发现 Linux 通过大 folio 减少了元数据工作,并将大部分冷读取复制与设备等待重叠,但这些机制依赖于访问建议、folio 粒度、缓存状态和后端执行。通过实验性内核原型,我们验证了乱序早期复制和机会性并行带来的额外机会,同时表明当 SSD 供应已足够时,更快的请求提交并不能改善端到端性能。为了解释设备等待,我们从有限的请求批次完成时间线中抽象出首次完成等待 $F$ 和后续完成容量 $B_{CQ}$。在真实 SSD 上的测量表征了它们如何随请求大小和批次大小变化,而 MQSim 实验将它们与内部设备机制联系起来,并区分了有限批次完成与持续吞吐量。然后,我们沿着实际的 folio--bio--request 映射构建了一个跨层等待链。独立校准的设备参数分别以不超过约 6.9% 和 3.2% 的误差预测了块层和 read_pages 处的等待。最后,四个用例将模型和等待链应用于并行提交、Linux 预读、依赖读布局以及轮询与睡眠。它们的收益、无收益和收益反转形成了从观察到建模再到解释和控制的闭环,为分层 I/O 栈中的性能决策提供了可测量的基础。

英文摘要

A simple read interface provides uniform functional semantics, but not equally simple or predictable performance behavior. We use synchronous large-buffer reads in Linux as an observation window and decompose the buffered-read path end to end, distinguishing work volume, processing time, and critical-path exposure. We find that Linux reduces metadata work through large folios and overlaps most cold-read copying with device waiting, but these mechanisms depend on access advice, folio granularity, cache state, and backend execution. Using an experimental kernel prototype, we validate additional opportunities from out-of-order early copying and opportunistic parallelism, while showing that faster request submission does not improve end-to-end performance when SSD supply is already sufficient. To explain device waiting, we abstract first-completion wait $F$ and subsequent completion capacity $B_{CQ}$ from finite request-batch completion timelines. Measurements on a real SSD characterize how they vary with request size and batch size, while MQSim experiments connect them to internal device mechanisms and distinguish finite-batch completion from sustained throughput. We then construct a cross-layer wait chain along the actual folio--bio--request mapping. Independently calibrated device parameters predict waiting at the block layer and at read_pages with errors no greater than approximately 6.9% and 3.2%, respectively. Finally, four use cases apply the model and wait chain to parallel submission, Linux readahead, dependent-read layout, and polling versus sleeping. Their gains, no gains, and gain reversals form a closed loop from observation and modeling to explanation and control, providing a measurable basis for performance decisions in layered I/O stacks.

Comments26 pages, 10 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑