用于大规模全切片图像嵌入提取的解耦式I/O主导型流水线
Decoupled I/O-Dominant Pipelines for Large-Scale Whole-Slide Image Embedding Extraction
- Oak Ridge National Laboratory(橡树岭国家实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大规模全切片图像嵌入提取的I/O与编排开销主导性能的问题,提出解耦I/O、计算与导入的流水线,实现高效高吞吐量的嵌入提取,将其转化为数据-centric系统问题。
AI中文摘要:
全切片图像(WSI)是计算病理学的核心,但尺寸极大,因此基于补丁的处理是基础模型推理的实用单元。然而在大规模场景下,快速生成和处理海量补丁会引入显著的I/O与编排开销,通常会主导端到端性能。本文提出一种用于大规模WSI嵌入提取的解耦式、I/O感知型流水线,将工作流分解为三个阶段:(1)补丁生成与暂存;(2)易并行嵌入推理;(3)分片向量数据库导入。该设计将数据移动与计算解耦,实现高效的补丁交付、可扩展的多节点推理且通信开销最小。所得系统生成一个分布式向量数据库,其中嵌入与丰富元数据(如患者、切片及补丁属性)持久关联,支持高效过滤、检索及下游复用。该表示数据库紧凑且可复用,适用于检索、分类、少样本学习等任务,尤其惠及低资源环境。研究表明,解耦I/O、计算与导入可实现大规模高吞吐量WSI嵌入提取;通过刻画扩展包络,证实中等并发以上存储占主导,将WSI嵌入提取重新定义为以数据为中心的系统问题,而非纯计算受限的工作负载。
英文摘要:
Whole-slide images (WSIs) are central to computational pathology but are prohibitively large, making patch-based processing the practical unit for foundation model inference. At scale, however, generating and handling massive numbers of patches on quickly introduces significant I/O and orchestration overhead, often dominating end-to-end performance. We present a decoupled, I/O-aware pipeline for large-scale WSI embedding extraction that decomposes the workflow into three stages: (1) patch generation and staging, (2) embarrassingly parallel embedding inference, and (3) sharded vector database ingestion. This design isolates data movement from compute, enabling efficient patch delivery, scalable multi-node inference with minimal communication. The resulting system produces a distributed vector database where embeddings are persistently coupled with rich metadata (e.g., patient, slide, and patch attributes), enabling efficient filtering, retrieval, and downstream reuse. This representation database is compact and reusable for tasks such as retrieval, classification, and few-shot learning, particularly benefiting low-resource environments. We show that decoupling I/O, computation, and ingestion enables high-throughput WSI embedding extraction at scale. By characterizing the scaling envelope, we demonstrate that storage dominates beyond moderate concurrency, reframing WSI embedding extraction as a data-centric systems problem rather than a purely compute-bound workload.