arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16953cs.AR

DTX:一种面向扩散模型与Transformer模型的以吞吐量优先的训练加速器

DTX: A Throughput-First Training Accelerator for Diffusion and Transformer Models

Shashank

首次发表
浏览论文内容

中文总结 AI 辅助

DTX是面向扩散与Transformer模型的吞吐量优先训练加速器,通过消除循环携带依赖、采用脉动阵列等设计实现高吞吐量,在多项测试中表现优异,流片后性能优于基线方案。

中文摘要 AI 辅助

DTX是一种面向扩散模型与Transformer模型的以吞吐量优先的训练加速器。任何通过单个FP32加法器序列化的求和操作都是循环携带的依赖关系,无论物理设计如何,都会将机器限制在接近2 FLOP/周期;DTX的设计使其不存在此类依赖链——每次归约操作都是流水线二叉树,每个FP算子都是启动间隔为1的两级流水线。一个8×8的权重固定脉动阵列,融合了偏置/激活/转换的收尾操作,搭配8通道向量单元、8通道融合AdamW流水线以及流水线Philox高斯源,由4槽VLIW指令在统一的64 KB tile空间中协同发射:达到216 FLOP/周期,约为每时钟循环携带最低值的108倍。由于无规范求和顺序,验证采用针对FP64黄金模型的容差方法,包含精确相等的特例处理,且具有可证明的严格边界(违反前提的程序测量值超出预算5340倍;17项测试、107108个元素,零失败)。语义门确认片上扩散-MLP运行将损失从56.4降至26.0;计数器级证明显示计算/DMA重叠维持峰值;分析性同节点分解将GPU的吞吐量每瓦比限定为6-10倍;sky130流片项目将脉动阵列硬化为布线后83.3 MHz的DRC兼容GDS——在同一节点和流程下,是优化的循环携带MAC基线的1.9倍。

英文摘要

DTX is a throughput-first training accelerator for diffusion and transformer models. Any summation serialized through a single FP32 adder is a loop-carried dependence that pins a machine near 2 FLOP/cycle regardless of physical design; DTX is built so no such chain exists anywhere -- every reduction is a pipelined binary tree, every FP operator a two-stage pipeline with initiation interval 1. An 8x8 weight-stationary systolic array with a fused bias/activation/cast epilogue, an 8-lane vector unit, an 8-lane fused AdamW pipeline, and a pipelined Philox Gaussian source are co-issued by a 4-slot VLIW word over a unified 64 KB tile space: 216 FLOP/cycle, roughly 108x the loop-carried floor per clock. With no canonical sum order, verification is tolerance-based against an FP64 golden model, with exact-equality carve-outs and a demonstrably tight bound (a premise-violating program measured 5,340x over budget; 17/17 tests, 107,108 elements, zero failures). Semantic gates confirm an on-device diffusion-MLP run reduces its loss (56.4 to 26.0), counter-level proof shows compute/DMA overlap sustains the peak, an analytical iso-node decomposition bounds the GPU comparison at 6-10x throughput per watt, and a sky130 campaign hardens the systolic array to DRC-clean GDS at 83.3 MHz post-route -- 1.9x an optimized loop-carried MAC baseline on the same node and flow.

↑