arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06801cs.CVcs.AI

MC-Sparse:解构并弥合扩散Transformer中的稠密-稀疏注意力差距

MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers

发表机构复旦大学 · 腾讯HY · 上海创新研究院
另 4 家 · 查看机构详情
  • Fudan University(复旦大学)
  • Tencent HY(腾讯HY)
  • Shanghai Innovation Institute(上海创新研究院)
  • MMLab, CUHK(香港中文大学多媒体实验室)
  • Nankai University(南开大学)
  • Peking University(北京大学)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiarui Chen, Zeqiang Lai, Jiangshan Wang, Ziheng Ouyang, Ye Huang, Xiangyu Yue, Cewu Lu, Chunchao Guo

首次发表
浏览论文内容

中文总结 AI 辅助

MC-Sparse通过元缓存和精确KV选择,在扩散Transformer中实现高效稀疏注意力,显著加速视频和3D生成,同时保持高保真度。

中文摘要 AI 辅助

稀疏注意力是降低扩散Transformer在长序列生成任务(如视频和高分辨率3D资产生成)中延迟的主要方法。然而,现有方法在高稀疏度下会降低生成质量和保真度。通过受控的oracle比较,我们将这种退化追溯到三个来源:token分组带来的约束、不准确的交互选择,以及token被丢弃时损失的注意力贡献。基于这一分析,我们提出了元缓存稀疏注意力(MC-Sparse),一种无需训练框架,它选择单个键值(KV)token,同时将相似查询组织成瓦片对齐的组,以实现高效的GPU执行。MC-Sparse缓存元数据,包括查询组、使用精确注意力概率选择的KV索引,以及稠密和稀疏注意力输出之间的残差,并在后续去噪步骤中重用它们。在视频和3D生成模型中,与现有稀疏注意力基线相比,MC-Sparse实现了更高的稠密注意力输出保真度和更大的去噪加速,且无明显质量退化。相对于稠密注意力,它在Minimax-H3-Base上实现了$1.80\times$的去噪加速,在3D资产生成上实现了$2.32\times$的加速,两者质量损失均可忽略。

英文摘要

Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided by this analysis, we propose Meta-Cached Sparse Attention (MC-Sparse), a training-free framework that selects individual key-value (KV) tokens while organizing similar queries into tile-aligned groups for efficient GPU execution. MC-Sparse caches metadata comprising query groups, KV indices selected using exact attention probabilities, and residuals between dense and sparse attention outputs, and reuses them across subsequent denoising steps. Across video and 3D generation models, MC-Sparse achieves higher fidelity to dense-attention outputs and larger denoising speedups than existing sparse-attention baselines, without visible quality degradation. Relative to dense attention, it delivers a $1.80\times$ denoising speedup on Minimax-H3-Base and a $2.32\times$ speedup on 3D asset generation, both with negligible quality loss.

补充信息

↑