世界模型硬件加速器
The World Model Hardware Accelerator
AI总结:
针对扩散变换器推理,提出延迟优先的WMHA加速器,利用编译期已知调度、权重固定点积阵列与单遍注意力流水线,在两种配置下实现23倍误差降低,零元素故障。
AI中文摘要:
扩散变换器反转了自回归解码所熟悉的算术方式。不存在逐令牌的循环:每个去噪步骤都是对静态形状的全序列前向传递,因此整个调度在编译时即可得知,唯一的串行维度是步骤数本身。我们在WMHA中利用了这种结构,这是一种延迟优先的扩散变换器推理加速器:一个超长指令字定序器从一条指令字发出四个引擎,一个权重固定的16x16双点积阵列流式传输FP8和BF16收缩,一个单遍在线softmax注意力流水线通过倾斜的软件流水线保持键和值驻留。该设计在冻结的微架构文档中规定,以可综合的SystemVerilog实现,并通过UVM环境对照双精度参考模型进行验证,其验收标准是语义性的:设备必须运行真实的去噪轨迹,并将相对于干净潜变量的均方误差至少降低十倍。它在两种综合配置下均实现了23倍的降低,在2.37亿个检查值中零元素故障。基于已发布模型形状构建的十一个应用基准(包括原始扩散变换器配置)在设备上运行,并报告了实测占用率以及单独标注的预测值。五个引擎在sky130中经过布线布局,具有寄生标注的时序和实测活动功耗;整个芯片已综合,阻止其布局布线的宿主限制与消除该限制的机器一起被量化。
英文摘要:
Diffusion transformers invert the arithmetic that autoregressive decoding made familiar. There is no token-by-token recurrence: every denoising step is a full-sequence forward pass over static shapes, so the entire schedule is known at compile time and the only serial dimension is the step count itself. We exploit that structure in WMHA, a latency-first diffusion-transformer inference accelerator: a very-long-instruction-word sequencer issues four engines from one instruction word, a weight-stationary 16x16 dual-dot array streams FP8 and BF16 contractions, and a single-pass online-softmax attention pipeline keeps keys and values resident through a skewed software pipeline. The design is specified in a frozen micro-architecture document, implemented in synthesizable SystemVerilog, and verified against a double-precision reference model by a UVM environment whose acceptance criterion is semantic: the device must run a real denoising trajectory and reduce mean squared error against a clean latent by at least a factor of ten. It does so by a factor of 23, at both synthesized configurations, with zero element failures across 237 million checked values. Eleven application benchmarks built from published model shapes, including the original diffusion-transformer configuration, run on the device and report measured occupancy beside separately labelled projections. Five engines are taken to routed layout in sky130 with parasitic-annotated timing and measured-activity power; the full chip is synthesized, and the host limit that stopped its place-and-route is quantified together with the machine that would remove it.