SAF3R:用于前馈3D重建Transformer的动态稀疏注意力
SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers
另 2 家 · 查看机构详情
- University of Pittsburgh(匹兹堡大学)
- University of Arizona(亚利桑那大学)
- Tongji University(同济大学)
- University of Central Florida(中佛罗里达大学)
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究前馈3D重建Transformer,分析其全局注意力模式,提出SAF3R框架,集成稀疏注意力机制、离线头部分析和在线适应策略,在保持质量同时提高稀疏率加速端到端速度。
中文摘要 AI 辅助
前馈3D重建(F3R)Transformer取得显著成功,但扩展到长图像序列有挑战。本文对多个F3R Transformer的全局注意力进行全面分析,提出SAF3R,将定制稀疏注意力机制与离线头部分析及在线适应策略集成,实验表明其在保持质量时实现高稀疏率和加速。
英文摘要
Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success. However, scaling them to long image sequences remains challenging, as the quadratic complexity of cross-view global attention quickly becomes the dominant computational bottleneck. While recent efforts attempt to improve efficiency through compressed or sparse attention, they fail to fully exploit the inherent sparsity and dynamic behavior of global attention. In this work, we present a comprehensive analysis of global attention across multiple F3R transformers and reveal that attention patterns are highly heterogeneous, dynamic, and extremely sparse across layers and attention heads. Motivated by these findings, we propose SAF3R, a training-free dynamic sparse attention framework tailored to F3R transformers. SAF3R integrates tailored sparse attention mechanisms with offline head profiling and an efficient online adaptation strategy to match input-dependent attention behaviors. Extensive experiments demonstrate that SAF3R achieves high sparsity ratios while preserving camera pose estimation and 3D reconstruction quality, translating into substantial end-to-end speedup on F3R transformers compared to existing methods. Code is available at https://github.com/jndeng/SAF3R