DanLing NestedTensor:面向深度学习的可组合多参差张量
DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning
- DanLing Team(丹灵团队)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出DanLing NestedTensor张量抽象,将多参差结构内建于张量本身,支持广播、变换和归约,在A100上实现2.74-3.39倍加速,峰值分配从38.08降至5.41 GiB。
AI中文摘要:
变长输入在深度学习中很常见,但稠密批处理会分配一个共享的外包络,并在填充上耗费计算。该成本在多个变化轴上成倍增加:显式配对状态分配$BN_{\max}^2$个位置,而非$\sum_i N_i^2$个。打包消除了这种浪费,但组合打包操作仍需要逻辑轴和样本边界,而扁平缓冲区不再暴露这些信息。我们提出DanLing NestedTensor,一种PyTorch张量抽象,使多参差结构成为张量本身的属性。打包值携带张量支撑的分区和逻辑维度顺序,因此广播创建参差轴,特征变换保留它们,归约消耗它们。相同的表示贯穿自动微分以及即时和编译执行。在A100上,与同模式填充相比,几何平均加速在四种BERT规模上即时执行为2.74倍,编译执行为3.39倍,在四种FCN骨干上即时执行为1.97倍。一个四块Pairformer风格工作负载在即时执行中,在方形长度范围内使用原生PyTorch内核比填充参考快2.40-4.32倍,其高变化批次中的峰值分配从38.08 GiB降至5.41 GiB。该张量接口使基于其支持的算子构建的模型代码能够组合高效的可变大小计算,而无需在任何调用点管理偏移。代码将在发表后公开发布。
英文摘要:
Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure a property of the tensor itself. Packed values carry tensor-backed partitions and logical dimension order, so broadcasting creates ragged axes, feature transformations retain them, and reductions consume them. The same representation carries through autograd and both eager and compiled execution. On an A100, the geometric-mean speedup over same-mode padding is 2.74$\times$ eager and 3.39$\times$ compiled across four BERT scales, and 1.97$\times$ eager across four FCN backbones. A four-block Pairformer-style workload runs 2.40-4.32$\times$ faster than a padded reference using native PyTorch kernels across square length regimes in eager execution, with peak allocation falling from 38.08 to 5.41 GiB on its high-variation batch. The tensor interface lets model code built from its supported operators compose efficient variable-size computation without managing offsets at any call site. Code will be released publicly upon publication.