基于可学习掩码注意力与支持令牌优化的无姿态前馈三维补全
Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement
浏览论文内容
中文总结 AI 辅助
FreeInpaint是一种无姿态前馈三维补全框架,通过可学习掩码注意力与支持令牌优化解决掩码输入适配问题,无需预计算相机姿态即可实现高质量三维场景补全,且推理速度快。
中文摘要 AI 辅助
三维场景补全旨在恢复编辑后三维场景中缺失或被遮挡的区域,同时保证几何与纹理一致性。然而现有方法通常需要精确校准的相机姿态,这限制了其在随意、野外场景中的适用性,并引入了额外的预处理开销。为克服这一局限,我们提出FreeInpaint,一种新颖的前馈框架,可直接从带有掩码区域的无姿态多视图图像生成完整且三维一致的场景。FreeInpaint的核心是扩展三维基础模型,将掩码区域从参考视图传播到其他无姿态视图,在保留模型恢复相机姿态与场景几何原生能力的同时,衔接三维重建与场景补全。我们的方法解决了将前馈三维基础模型适配掩码输入时的两个关键挑战:其一,掩码区域会破坏跨视图对应推理,降低姿态估计与几何恢复性能,为此我们引入可学习掩码注意力机制,该机制在保留可靠观测空间锚定的同时,允许掩码区域在更深层逐步吸收有用上下文;其二,在严重遮挡下,单次前向传播往往缺乏足够外观证据以实现高保真补全,因此我们提出支持令牌优化策略,该策略将扩散生成的支持证据作为置信度加权辅助令牌注入,以在保留原始空间锚定的同时优化观测不足区域。在不同数据集上开展的大量实验表明,FreeInpaint实现了更优的补全质量,消除了对预计算相机姿态的依赖,同时保持了快速推理速度。项目页面为this https URL。
英文摘要
3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency. Existing approaches, however, typically require accurately calibrated camera poses, which restricts their applicability in casual, in-the-wild scenarios and introduces additional preprocessing overhead. To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions. At its core, FreeInpaint extends a 3D foundation model to propagate masked regions from a reference view to other unposed views, bridging 3D reconstruction and scene inpainting while preserving the model's native ability to recover camera poses and scene geometry. Our method addresses two key challenges in adapting feed-forward 3D foundation models to masked inputs. First, masked regions can corrupt cross-view correspondence reasoning, degrading pose estimation and geometry recovery. To address this, we introduce a Learnable Mask Attention mechanism that preserves the spatial anchoring of reliable observations while allowing masked regions to progressively absorb useful context in deeper layers. Second, under severe occlusions, a single forward pass often lacks sufficient appearance evidence for high-fidelity completion. Therefore, we propose a Support Token Refinement strategy, which injects diffusion-generated support evidence as confidence-weighted auxiliary tokens to refine under-observed regions while preserving the original spatial anchor. Extensive experiments across diverse datasets demonstrate that FreeInpaint achieves superior inpainting quality, eliminating the reliance on pre-computed camera poses while keeping a fast inference speed. The project page is https://rorisis.github.io/FreeInpaint/.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。