ReSS:用于3D视觉Transformer的残差恢复稀疏注意力
ReSS: Residual-Restoring Sparse Attention for 3D Vision Transformers
浏览论文内容
中文总结 AI 辅助
针对3D视觉Transformer全局注意力计算成本高的问题,提出ReSS方法,将块选择从最大化注意力质量转为最小化残差流漂移,通过迭代残差恢复优化丢弃集,在多个骨干和数据集上优于现有方法。
中文摘要 AI 辅助
诸如VGGT之类的3D视觉Transformer在单次前向传播中从多视图图像预测相机姿态和场景几何,但随着视图数量的增加,对所有拼接的视图标记进行的全局注意力在计算中占主导地位。为降低这一成本,SparseVGGT和HeSS在块级别稀疏化注意力,两者都保留具有高注意力概率的块。然而,我们观察到注意力概率并不能很好地预测当某个块被移除时模型行为实际改变的程度,并且我们表明这种不匹配正是随着稀疏度增加性能崩溃的原因。在本文中,我们提出ReSS(残差恢复稀疏注意力),它将块选择从最大化保留注意力质量的问题重新定义为最小化稀疏化在残差流中留下的漂移的问题。我们引入了一个漂移分数来量化每个块对残差的偏移程度,并且由于丢弃集的漂移取决于贡献向量的方向而不仅仅是其幅度,我们提出了一种迭代残差恢复程序,以整体方式优化丢弃集。在三个骨干网络和五个数据集上,ReSS在匹配的稀疏度下比先前方法更好地保持了密集性能。另外两个结果支持漂移是控制稀疏化成本的量:最大化漂移比随机选择更快地降低性能,并且相对于实际漂移而非稀疏度绘制时,所有方法近似落在一条曲线上。代码可在以下https URL获取。
英文摘要
3D vision transformers such as VGGT predict camera poses and scene geometry from multi-view images in a single forward pass, but their global attention over all concatenated view tokens dominates computation as the number of views grows. To reduce this cost, SparseVGGT and HeSS sparsify attention at the block level, and both retain blocks with high attention probability. However, we observe that attention probability poorly predicts how much the model's behavior actually changes when a block is removed, and we show that this mismatch is why performance collapses as sparsity increases. In this paper, we propose ReSS (ReSidual-ReStoring Sparse Attention), which recasts block selection from a problem of maximizing the retained attention mass to one of minimizing the drift that sparsification leaves in the residual stream. We introduce a drift score that quantifies how much each block shifts the residual, and, since the drift of a drop set depends on the directions of the contribution vectors rather than on their magnitudes alone, an iterative residual restoration procedure that refines the drop set as a whole. Across three backbones and five datasets, ReSS preserves dense performance better than prior methods at matched sparsity. Two further results support drift as the quantity that governs the cost of sparsification: maximizing drift degrades performance faster than random selection, and plotted against realized drift instead of sparsity, all methods fall approximately onto a single curve. Code is available at https://github.com/libary753/ReSS.
发表机构
- Seoul National University(首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。