AI 中文总结
DraftAttention2提出无需训练的框架,利用低分辨率草稿注意力图联合选择注意力块并分配混合精度,在少步视频扩散中实现质量与效率的优越权衡。
AI 中文摘要
视频生成在内容创作和娱乐领域具有广泛应用。扩散变换器提升了生成视频的质量,但随着视频分辨率和时长的增加,对时空令牌的注意力计算成本日益高昂。我们提出了DraftAttention2,一个无需训练的框架,利用低分辨率草稿注意力图来联合选择注意力块并分配其数值精度。具体而言,空间二维平均池化和最大池化的查询与键捕获互补的区域统计信息以估计块的重要性,共享排序为重要块分配更高精度,为较不重要的保留块分配较低精度,并在可配置预算下跳过其余块。我们的分析将稀疏化误差与注意力加权量化误差分离,确立了何时用低位计算恢复被跳过的交互能收紧输出误差界。该分析促使在低精度下保留更多交互,同时为注意力质量更大的块保留更高精度。为了将这些细粒度分配转化为实际加速,我们进一步开发了融合操作数准备和具有精度特定阶段的单一注意力内核,跨精度共享数据移动、softmax统计和输出累积。实验表明,我们的方法在质量-效率权衡上优于现有的高效视频生成方法。值得注意的是,其优势在少步视频扩散中尤为显著,其中将稀疏性与4位和8位混合精度计算联合结合,在保持显著加速的同时大幅提升生成质量。代码可在该https URL获取。
英文摘要
Video generation has broad applications in content creation and entertainment. Diffusion transformers have advanced the quality of generated videos, but attention over spatiotemporal tokens becomes increasingly expensive as video resolution and duration increase. We present DraftAttention2, a training-free framework that uses the low-resolution draft attention map to jointly select attention blocks and assign their numerical precision. Specifically, spatial 2D average- and max-pooled queries and keys capture complementary regional statistics to estimate block importance, and a shared ranking assigns higher precision to important blocks, lower precision to less important retained blocks, and skips the rest under configurable budgets. Our analysis separates sparsification error from attention-weighted quantization error, establishing when recovering skipped interactions with low-bit computation tightens the output-error bound. This analysis motivates retaining more interactions at low precision while reserving higher precision for blocks with larger attention mass. To translate these fine-grained assignments into practical speedups, we further develop fused operand preparation and a single attention kernel with precision-specific phases, sharing data movement, softmax statistics, and output accumulation across precisions. Experiments demonstrate that our method achieves a superior quality-efficiency trade-off over existing efficient video generation methods. Notably, its advantage is particularly pronounced for few-step video diffusion, where jointly combining sparsity with 4- and 8-bit mixed-precision computation substantially improves generation quality while retaining significant acceleration. Code is available at https://github.com/anemoi-project/anemoi
CommentsPreprint Version