arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Zellige:用于混合图像-视频DiT训练的可塑序列放置

Zellige: Moldable Sequence Placement for Mixed Image-Video DiT Training

Guangyu Xiang, Xueze Kang, Minwei Zhao, Yuxin Wang, Shaohuai Shi, Lin Zhang, Xiaowen Chu

arXiv 2608.01150首次发表:更新:

AI 中文总结

Zellige是一种可塑序列放置系统,通过硬件分析器、两阶段规划器和合并注意力引擎解决混合图像-视频DiT训练的GPU序列放置问题,在多GPU环境下性能优于KnapFormer。

AI 中文摘要

高质量视频生成需要在图像和视频数据上联合训练扩散变换器(Diffusion Transformers,DiTs),这在GPU间带来了混合长度序列的训练问题。现有系统依赖数据并行(DP)、上下文并行(CP)或其组合;我们将这些设计建模为不相交组放置,并证明它们面临组间负载不平衡与组内通信冗余之间的根本权衡。我们提出Zellige,一种可塑序列放置系统,可联合选择每个序列的并行配置和参与秩。Zellige包含三个组件:硬件分析器,用于估计候选放置的执行时间和内存消耗;两阶段规划器,用于平衡计算密集型锚序列并将较轻的填充序列打包到剩余容量中;以及合并注意力引擎,可高效执行完整序列和分布式注意力分片。在21个规划方案中,硬件分析器预测步长总耗时和峰值已分配内存的平均绝对百分比误差分别为3.4%和1.5%。两阶段规划器处理每个批次耗时33-119毫秒,显著快于联合优化所有序列的参考放置方法,而它们建模的总耗时差异最多为0.32%。在端到端评估中,Zellige在16个A800 GPU上比KnapFormer快1.12-1.48倍,在32个A6000 GPU上快1.27-1.54倍。

英文摘要

High-quality video generation requires training Diffusion Transformers (DiTs) jointly on image and video data, posing a mixed-length sequence training problem across GPUs. Existing systems rely on data parallelism (DP), context parallelism (CP), or their combination; we model these designs as disjoint-group placement and prove that they face a fundamental tradeoff between inter-group load imbalance and intra-group communication redundancy. We present Zellige, a moldable sequence placement system that jointly selects each sequence's parallelism configuration and participating ranks. Zellige consists of three components: a hardware profiler that estimates the execution time and memory consumption of candidate placements, a two-stage planner that balances compute-heavy anchor sequences and packs lighter filler sequences into the remaining capacity, and a coalesced attention engine that efficiently executes whole sequences alongside distributed-attention shards. Across 21 plans, the hardware profile predicts step makespan and peak allocated memory with mean absolute percentage errors of $3.4%$ and $1.5%$, respectively. The two-stage planner solves each batch in 33--119 ms, significantly faster than a joint-placement reference that optimizes all sequences together, while their modeled makespans differ by at most $0.32%$. In end-to-end evaluations, Zellige outperforms KnapFormer by $1.12$--$1.48\times$ on 16 A800 GPUs and $1.27$--$1.54\times$ on 32 A6000 GPUs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑