arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更少上下文,更好几何:用于鲁棒3D基础模型的掩码几何编码器

Less Context, Better Geometry: Masked Geometric Encoder for Robust 3D Foundation Models

Zhimin Shao, Xijun Liu, Zhaoliang Zhang, Yutao Tang, Abhay Yadav, Rama Chellappa, Cheng Peng

arXiv 2610.06813首次发表:更新:

发表机构

Johns Hopkins University; University of Virginia(约翰霍普金斯大学; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出掩码几何编码器(MGE),通过训练时丢弃帧令牌和教师蒸馏学习鲁棒几何表示,并配合锚点引导自适应令牌合并,在遮挡和相似视图下提升3D重建性能与推理效率。

AI 中文摘要

3D基础模型的最新进展通过利用从大量空间数据中学习到的3D先验,实现了快速的3D重建和相机标定。然而,全对全的全局注意力设计导致二次复杂度,并限制了长序列推理;无约束的跨视图交互也可能传播来自遮挡或视觉上相似但几何上遥远视图的不可靠证据。在本文中,我们引入了一种掩码几何编码器(MGE),它促进了在不完整跨视图上下文下学习鲁棒的几何表示。在训练期间,MGE策略性地从全局注意力中丢弃帧令牌,并从预训练的全上下文教师模型中蒸馏。这使得模型能够学习本质上更丰富的每帧表示,同时提供足够的中间监督以避免性能下降。通过大量实验,我们表明MGE在遮挡和相似视图(doppelganger views)下带来了更强的性能,同时在标准基准上保持了高性能。这种更丰富的帧表示也导致了推理期间更有效的令牌减少。为此,我们开发了一种新颖的锚点引导自适应令牌合并技术,该技术保留代表性的锚点帧,同时联合合并来自其余视图的冗余令牌。与其他高效推理方法相比,我们能够在实现推理加速的同时,持续保持更高的重建质量,特别是在有限视图设置中。

英文摘要

Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions also can propagate unreliable evidence from occluded or visually similar but geometrically distant views. In this paper, We introduce a Masked Geometric Encoder (MGE), which promotes the learning of robust geometric representations under incomplete cross-view context. During training, MGE strategically drops frame tokens from global attention and distills from a pretrained full-context teacher model. This allows the model to learn an intrinsically richer per-frame representation while providing sufficient intermediate supervision to avoid performance degradation. Through extensive experiments, we show that MGE leads to much stronger performance under occlusion and doppelganger views while retaining high performance on standard benchmarks. Such a richer frame representation also leads to more effective token reduction during inference. To this end, we develop a novel Anchor-Guided Adaptive token merging technique that preserves representative anchor frames while jointly merging redundant tokens from the remaining views. Compared to other efficient inference approaches, we can achieve inference speedup while consistently maintaining higher reconstruction quality, particularly in limited-view settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑