arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12780cs.CV

SCOPE:用于稀疏视频注意力的在线分头Top-K估计的子空间聚类

SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention

Qi Zhao, Qirui Li, Hanlin Tang, Yiduo Li, Zhen Guo, Cuifeng Shen, Chao Xu, Zhaosheng Chi, Xiaojin Lu, Kan Liu, Tao Lan, Lin Qu, Xi Li

首次发表
浏览论文内容

中文总结 AI 辅助

SCOPE是一种无训练稀疏注意力框架,通过子空间聚类与在线分头Top-k估计解决视频DiTs的高成本问题,在6种配置中优于基线,实现1.99倍加速且保持28.46 dB PSNR。

中文摘要 AI 辅助

扩散Transformer(DiTs)在时空令牌上会产生二次自注意力成本。现有的无训练稀疏注意力方法通常从块级或聚类级代理分数构建稀疏掩码,这会掩盖键之间的细粒度差异,并在激进稀疏性下遗漏高贡献键。此外,此类代理分数可能产生过于集中的softmax分布,导致Top-p为某些查询聚类保留过少的键。尽管固定的Top-k最小值可缓解这种失败模式,但共享值无法适应不同头和输入的变化。为解决这两个限制,我们提出SCOPE,一种无训练稀疏注意力框架,结合3D-RoPE对齐的键子空间聚类与在线分头Top-k估计,用于高效视频DiT推理。SCOPE将RoPE后的键划分为时间、高度和宽度子空间,独立对其聚类,并通过查找表聚合对应的质心分数,以获取每个查询聚类的每个键的代理分数。在现有混合Top-p/固定Top-k选择的基础上,SCOPE通过按查询聚类大小加权平均每个头内初始保留的键数,在线推导头特定的Top-k值,并仅为初始保留键数低于该值的查询聚类选择额外的键。随后对选定的原始键和值计算稀疏注意力。在6种模型-任务配置中,SCOPE在保真度和延迟上始终优于现有无训练基线,在720p HunyuanVideo上相对于密集注意力实现了高达1.99倍的端到端加速,且PSNR为28.46 dB。

英文摘要

Diffusion Transformers (DiTs) incur quadratic self-attention cost over spatiotemporal tokens. Existing training-free sparse attention methods often construct sparse masks from block-level or cluster-level proxy scores, which can obscure fine-grained differences among keys and miss high contribution keys under aggressive sparsity. Moreover, such proxy scores may yield overly concentrated softmax distributions, causing Top-$p$ to retain too few keys for some query clusters. Although a fixed Top-$k$ minimum alleviates this failure mode, a shared value cannot adapt to variations across heads and inputs. To address both limitations, we propose SCOPE, a training-free sparse attention framework that combines 3D-RoPE-aligned key subspace clustering with online per-head Top-$k$ estimation for efficient video-DiT inference. SCOPE partitions post-RoPE keys into temporal, height, and width subspaces, clusters them independently, and aggregates the corresponding centroid scores through lookup tables to obtain per key proxy scores for each query cluster. Building on existing hybrid Top-$p$/fixed Top-$k$ selection, SCOPE derives a head-specific Top-$k$ value online by averaging the initial retained key counts within each head, weighted by query cluster size, and selects additional keys only for query clusters whose initial retained key counts fall below this value. Sparse attention is then computed over the selected original keys and values. Across six model--task configurations, SCOPE consistently outperforms existing training-free baselines in both fidelity and latency, achieving up to a $1.99\times$ end-to-end speedup on 720p HunyuanVideo with $28.46$ dB PSNR relative to dense attention.

发表机构

  • Zhejiang University(浙江大学)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

↑