arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RayOrch:为基座模型数据准备编程与执行谱系控制的多粒度数据流

RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation

Xiaochen Ma, Zimo Meng, Junzhu Liang, Youhe Jiang, Yue Cheng, Hao Liang, Bohan Zeng, Dengchun Li, Lu Ma, Zhengyang Zhao, Zhen Hao Wong, Runming He, Meiyi Qiang, Jiangtao Guan, Binhang Yuan, Wentao Zhang

arXiv 2609.18703首次发表:更新:

AI 中文总结

RayOrch提出一种编程模型与分布式执行引擎,通过声明式扩展和收集操作保留父子关系,实现多粒度数据流的高效批处理,在GPU上显著加速基座模型数据准备流水线。

AI 中文摘要

为基座模型准备高质量训练数据需要可扩展的流水线,将异构文档和视频转换为结构化记录。此类流水线将每个父项扩展为有序且依赖于输入的子项序列,其数量可能呈长尾分布。GPU应在保留父项关系、子项顺序、完成状态和结果路由的同时,跨父项进行子项批处理。现有系统要么将并行性隐藏在粗粒度作业之后,要么暴露扁平记录,迫使应用程序自行管理谱系和重新分组。我们提出RayOrch,一种编程模型和分布式执行引擎,在整个执行过程中保留父子关系。程序声明有序的可变基数扩展和匹配的收集操作。编译器验证每一对,而运行时记录子项成员资格、直接父项、不可变序号和终止状态。每调用FIFO就绪队列跨父项批处理就绪子项。收集操作根据声明的成员资格和序号重建结果,而非批处理边界或完成顺序。一旦所有必需子项变为终止状态,父项即可推进。类型化的父项作用域失败抑制失败父项的未调度兄弟项,同时允许不相关的父项继续。在NVIDIA H20 GPU上,RayOrch在将MinerU从4个GPU扩展到64个GPU时实现了15.14倍加速,在将视频流水线从8个GPU扩展到64个GPU时实现了7.82倍加速。与Ray Data相比,在MinerU上将端到端时间减少了13.1%,与Daft相比减少了29.0%,在Docling上与Ray Data相比减少了16.0%。代码可在https://this URL获取。

英文摘要

Preparing high quality training data for foundation models requires scalable pipelines that transform heterogeneous documents and videos into structured records. Such pipelines expand each parent item into an ordered and input dependent sequence of children, whose counts may be long tailed. GPUs should batch children across parents while preserving parent relationships, child order, completion status, and result routing. Existing systems either hide parallelism behind coarse grained jobs or expose flat records that force applications to manage lineage and regrouping. We present RayOrch, a programming model and distributed execution engine that preserves parent child relations throughout execution. Programs declare ordered variable cardinality expansions and matching gathers. The compiler validates each pair, while the runtime records child membership, immediate parents, immutable ordinals, and terminal states. Per Call FIFO Ready Queues batch ready children across parents. Gathers reconstruct results from declared membership and ordinals rather than batch boundaries or completion order. Parents can advance as soon as all required children become terminal. Typed parent scoped failures suppress undispatched siblings of the failed parent while allowing unrelated parents to continue. On NVIDIA H20 GPUs, RayOrch achieves 15.14 times speedup when scaling MinerU from 4 to 64 GPUs and 7.82 times speedup when scaling a video pipeline from 8 to 64 GPUs. It reduces end to end time by 13.1 percent versus Ray Data and 29.0 percent versus Daft on MinerU, and by 16.0 percent versus Ray Data on Docling. Code available at https://github.com/OpenDCAI/RayOrch .

CommentsTechnical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑