VASC:面向高效3D重建的值感知稀疏注意力与跨层记忆
VASC: Value-Aware Sparse Attention with Cross-Layer Memory for Efficient 3D Reconstruction
浏览论文内容
中文总结 AI 辅助
提出免训练的VASC稀疏注意力方法,通过值感知块选择与跨层记忆,在固定预算下提升3D重建的位姿估计质量,推理加速达2.29倍。
中文摘要 AI 辅助
前馈式3D视觉模型(如VGGT)在单次前向传播中统一了相机估计和密集场景重建,取得了显著进展。然而,其二次复杂度的全局注意力使得长图像序列的计算代价高昂,而现有的稀疏方法可能偏向于高注意力但值冗余的区域。为解决这些限制,我们提出了VASC,一种免训练的稀疏注意力方法,结合了值感知块选择和执行感知的跨层记忆。我们的值感知块选择整合了池化查询-键相关性与邻近值对比,在保留查询相关和独特内容的同时减少冗余。跨层记忆跨层追踪未满足的需求,并根据实际执行更新该状态,使得先前未被充分服务的块在固定计算预算下能够竞争。在7Scenes和NeuralRGB-D数据集上,使用VGGT和π³进行的实验表明,与FasterVGGT相比,位姿估计和重建质量均有提升,同时推理速度比密集VGGT快达2.29倍。代码可在该https URL获取。
英文摘要
Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. To address these limitations, we introduce VASC, a training-free sparse attention method combining value-aware block selection and execution-aware cross-layer memory. Our value-aware block selection integrates pooled query--key relevance with neighboring value contrast, reducing redundancy while preserving query-relevant and distinctive content. Cross-layer memory tracks unserved demand across layers and updates this state according to actual execution, enabling previously underserved blocks to compete under a fixed computation budget. Experiments on 7Scenes and NeuralRGB-D with VGGT and $π^3$ demonstrate improved pose estimation and reconstruction quality compared with FasterVGGT, together with up to $2.29\times$ faster inference than dense VGGT. Code is available at https://github.com/kosakayamahoo-design/VASC.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。