FoldPipe:基于异步预取的原生分子分片的有界远程流式传输
FoldPipe: Bounded Remote Streaming of Native Molecular Shards with Asynchronous Prefetch
AI总结:
FoldPipe是针对原生.pt分子分片的轻量级Python编排层,通过异步预取实现I/O-计算重叠,在MD17阿司匹林的SchNet实验中验证了重叠机制,但未明确可靠的 wall-clock 速度优势。
AI中文摘要:
在短暂或内存受限的加速器实例上训练分子机器学习模型,需要反复从远程存储中检索预处理的分子图。FoldPipe是一个轻量级Python编排层,用于已分片的PyTorch和PyTorch Geometric数据;它在消费端训练当前分片时,会在后台线程中提前检索一个分片,同时相对于总数据集大小保持活跃分片有效载荷的数量有界。异步预取和有界缓冲是成熟的系统技术,而非新颖的调度算法。FoldPipe的贡献在于针对原生.pt分子分片的小型集成,以及对其运行模式的源固定经验表征。我们在Tesla T4上对MD17阿司匹林的SchNet能量-力工作负载进行评估,采用20次配对、顺序交替的基准测试轮次,每轮处理5个固定分片,包含25000个结构。FoldPipe记录的平均I/O-计算重叠时间为16.33秒,而顺序有界基准的该值为0;FoldPipe的平均轮次时间为76.78秒,基准为83.37秒,但几何平均配对加速比为1.059倍,95%自举区间为0.878倍至1.288倍。因此,实验验证了重叠机制,但在观测到的公共网络变异性下,关于可靠的 wall-clock 速度优势尚无定论。
英文摘要:
Training molecular machine-learning models on ephemeral or memory-constrained accelerator instances can require repeatedly retrieving preprocessed molecular graphs from remote storage. FoldPipe is a lightweight Python orchestration layer for already-sharded PyTorch and PyTorch Geometric data. It retrieves one shard ahead in a background thread while the consumer trains on the current shard, keeping the number of live shard payloads bounded with respect to total dataset size. Asynchronous prefetch and bounded buffering are established systems techniques rather than novel scheduling algorithms. FoldPipe's contribution is a small integration targeted at native .pt molecular shards together with a source-pinned empirical characterization of its operating regime. We evaluate a SchNet energy-and-force workload on MD17 aspirin using 20 paired, order-alternating benchmark passes on a Tesla T4. Each pass processes five pinned shards containing 25,000 structures. FoldPipe records 16.33 s mean I/O-compute overlap, compared with zero by construction for the sequential bounded baseline. Mean pass time is 76.78 s for FoldPipe and 83.37 s for the baseline. However, the geometric mean paired speedup is $1.059\times$ with a 95% bootstrap interval from $0.878\times$ to $1.288\times$. The experiment therefore verifies the overlap mechanism but is inconclusive about a reliable wall-clock speed advantage under the observed public-network variability.