arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编辑前先观察:用于3D高斯点渲染编辑的注意力引导相机放置和多视图对齐

Look Before You Edit: Attention-Guided Camera Placement and Multi-View Alignment for 3D Gaussian Splatting Editing

Jaeyeon Park, Taeho Kang, Youngki Lee

arXiv 2607.19777首次发表:更新:

发表机构

Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对3D高斯点渲染编辑受固定相机限制的问题,提出LB-Edit框架。通过注意力引导相机放置确定编辑相机位置,多视图注意力对齐使视图编辑一致。实验表明该方法在多方面表现优,减少视图数量并降低延迟。

AI 中文摘要

基于文本的3D高斯点渲染(3DGS)场景编辑通常将二维扩散编辑器应用于从固定训练相机渲染的视图,限制了编辑的空间覆盖范围以及用户在复杂场景中针对特定对象的自由度。我们提出了LB-Edit框架,它解决了两个相关问题:为局部编辑放置编辑相机的位置,以及如何使每个视图的编辑相互一致,以便在微调后3D场景保持一致。首先,注意力引导编辑相机放置(ACP)在多个候选相机距离处探测扩散模型的自注意力和交叉注意力,以找到注意力在感兴趣区域内良好包含的位置,然后在该注意力最优距离处放置一组紧凑、几何上多样的编辑相机。其次,多视图注意力对齐(MAA)沿着两个轴引导编辑器在不同视图间进行相同的编辑:它通过令牌级对应共享自注意力特征来对齐外观,并将交叉注意力图提升到3D高斯上作为共享的3D注意力场来对齐空间位置,抑制外观和空间漂移。在多对象和单对象场景上的实验表明,我们的方法在指令保真度、多视图一致性和编辑局部性方面实现了最高的用户偏好,使用少至5个编辑视图,并且比现有方法减少延迟高达7倍。

英文摘要

Text-driven 3D scene editing with 3D Gaussian Splatting (3DGS) typically applies a 2D diffusion editor to views rendered from fixed training cameras, limiting both the spatial coverage of edits and the user's freedom to target specific objects in complex scenes. We present LB-Edit, a framework that addresses two coupled problems: where to place editing cameras for localized edits, and how to make per-view edits agree with one another so that the 3D scene remains consistent after fine-tuning. First, Attention-Guided Editing Camera Placement (ACP) probes the diffusion model's self- and cross-attention at multiple candidate camera distances to find where attention is well-contained in the region of interest, then places a compact, geometrically diverse editing camera set at that attention-optimal distance. Second, Multi-view Attention Alignment (MAA) steers the editor toward the same edit across views along two axes: it aligns appearance by sharing self-attention features via token-level correspondence, and aligns spatial location by lifting cross-attention maps onto the 3D Gaussians as a shared 3D attention field, suppressing both appearance and spatial drift. Experiments on multi-object and single-object scenes show that our method achieves the highest user preference in instruction fidelity, multi-view consistency, and editing locality, using as few as 5 editing views and reducing latency by up to 7x over existing methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑