arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21438cs.CVcs.MA

DesignAgent3D:通过类设计师多模态推理实现交互式3D场景编辑

DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning

  • University of Michigan(密歇根大学)
  • University of Notre Dame(圣母大学)
  • Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiujin Liu, Tianyu Yang, Yilun Zhao, Xiangliang Zhang

AI总结:

DesignAgent3D是一种交互式多模态智能体框架,通过类设计师的规划-感知-行动范式解决文本引导3D场景编辑的缺陷,在NeRF和3D Gaussian Splatting上均优于现有方法,实现了更优的语义对齐、空间定位精度和多视图一致性。

AI中文摘要:

文本引导的3D场景编辑为修改重建环境提供了直观界面,但仍存在挑战:自然语言设计请求常语义不明确,且需在杂乱3D场景中进行定位。现有方法通常将该任务表述为基于单一提示的一次性条件生成,无法解决用户意图模糊或实现精确空间定位,因此存在严重的对象定位漂移、遮挡下跟踪失败以及臭名昭著的多视图“贴纸效应”。为克服这些局限,本文提出DesignAgent3D,这是一种交互式多模态智能体框架,将3D场景编辑重新表述为类设计师的规划-感知-行动范式。该智能体首先通过与用户交互规划,澄清不明确的设计目标;接着通过将预期编辑定位到3D场景中的特定对象或区域进行感知;最后通过应用可控视觉修改同时保持场景一致性来执行操作。这些编辑还会被整合到底层3D表示中,支持持久且多视图一致的新视图渲染。在NeRF和3D Gaussian Splatting两种主干上进行的大量实验表明,DesignAgent3D显著优于最先进的基线方法,实现了更优的语义意图对齐、无可挑剔的空间定位精度以及高保真的多视图一致性。

英文摘要:

Text guided 3D scene editing provides an intuitive interface for modifying reconstructed environments, but remains difficult because natural language design requests are often semantically underspecified and must be grounded in cluttered 3D scenes. Existing methods typically formulate the task as one-shot conditional generation from a single prompt, failing to resolve ambiguous user intents or achieve precise spatial grounding. Consequently, they suffer from severe object localization drift, tracking failure under occlusions, and the notorious multi-view "sticker effect." To overcome these limitations, we present DesignAgent3D, an interactive multimodal agentic framework that reformulates 3D scene editing as a designer-like Plan-Perceive-Act paradigm. The agent first plans by interacting with the user to clarify underspecified design goals, then perceives by grounding the intended edit to specific objects or regions in the 3D scene, and finally acts by applying controlled visual modifications while preserving scene consistency. The edits are further integrated into the underlying 3D representation, supporting persistent and multi-view consistent novel-view rendering. Extensive experiments across both NeRF and 3D Gaussian Splatting backbones demonstrate that DesignAgent3D significantly outperforms state-of-the-art baselines, delivering superior semantic intent alignment, impeccable spatial localization accuracy, and high-fidelity multi-view consistency.

↑