发表机构
Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SparsePR算法,结合响应耦合划分与探针拟合残差重建,在四个视频生成和世界模型上实现注意力重建误差降低,22.0%-26.0%执行对密度下保持生成质量,获1.48-2.61倍端到端加速。
AI 中文摘要
无训练的块稀疏注意力可加速视频Transformer,但逐行注意力集中本身并不能指定可执行的稀疏算子。共享块路由的查询可能具有重叠度低的支撑集,且仅保留的注意力质量无法确定被跳过交互产生的后Softmax误差。本文表明,划分几何同时影响池化支撑集和稀疏输出的剩余残差的可预测性。我们引入SparsePR,其结合了响应耦合划分与探针拟合残差重建。采样查询的键响应形成配对键/值组,其质心诱导用于共享路由的查询-响应坐标。一小部分精确查询行随后在探针残差观察到的输出子空间内,从稀疏输出校准特定调用的仿射校正。在四个异构视频生成和世界模型上,SparsePR持续降低注意力重建误差。消融实验显示,探针拟合占该降低的大部分,而响应耦合划分降低硬丢弃误差并在有限探针预算下提升重建效果。SparsePR在22.0%-26.0%的已实现执行对密度下保持生成质量,同时实现1.48倍至2.61倍的端到端加速。项目页面:this https URL
英文摘要
Block-sparse attention accelerates video transformers by grouping queries and keys and computing attention only between selected groups. However, it faces two challenges: (1) queries sharing a block route may attend to different keys; and (2) retaining most attention mass does not necessarily preserve the attention output. We find that partition choice affects both the overlap among grouped queries' preferred supports and how well an affine function of the sparse output can represent the difference between dense and sparse outputs. We introduce SparsePR, a training-free method combining Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. We first group paired key/value tokens by how their keys respond to sampled queries. We then group queries by their responses to the resulting key-group centroids for shared sparse routing. We compute exact attention on a small set of query rows and fit an affine correction to the differences between their dense and sparse outputs using weighted ridge regression. The fitted map predicts a residual for each unprobed query, which we add to its sparse output. We measure attention-output error as the relative L2 error with respect to each query's dense output. Across four video generation and world models, SparsePR reduces this error, averaged over unprobed queries, by 61.6--90.5% compared with semantic block-sparse attention without residual correction, at 22% executed-pair density including exact probes. SparsePR achieves 1.48x--2.61x end-to-end speedups over dense attention, with dense-reference PSNR of 24.42--31.84 dB at 21.9--26.0% realized executed-pair density. Project page: https://pardistaghavi.github.io/SparsePR-website/
Comments26 pages, 10 figures. Project page: https://pardistaghavi.github.io/SparsePR-website/