arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.14168cs.CV

MUSE: 基于记忆的增量需求满足的智能体3D场景创作

MUSE: Agentic 3D Scene Authoring via Memory-Grounded Incremental Requirement Satisfaction

  • East China Normal University(华东师范大学)
  • Fudan University(复旦大学)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Ruijie Xu, Xinnan Zhu, Jiayu Ying, Daoguo Dong, Yuzhou Ji, Xin Tan

AI总结:

提出MUSE多智能体框架,通过增量需求满足实现可控3D场景构建与编辑,在AuthorBench上显著提升目标达成率和场景保持率。

AI中文摘要:

文本驱动的3D场景生成是数字内容创作、具身AI仿真和交互设计中的一项有前景的技术,然而实际工作流程通常需要在保留非目标内容的同时对现有场景进行细化、扩展或修正。现有方法可以生成逼真且结构合理的场景,但它们通常缺乏具有需求级状态跟踪的可编辑性,因此部分级故障常常导致全场景重新生成或人工干预。为应对这一挑战,我们将可控3D场景创作形式化为增量需求满足,统一了构建和编辑。在本文中,我们提出了MUSE,一个基于记忆的多智能体框架,其中架构师将指令编译为结构化需求,雕刻师执行局部场景操作,检查员验证每一步并更新工作记忆、场景记忆和技能记忆。为了评估需求级可控性和保留感知编辑,我们引入了AuthorBench,提供145个受限构建案例和1584个保留感知编辑池,并配有外部结构化检查。在全构建案例上,MUSE将全目标成功率从37.9提升至80.7,表面约束满足率从35.0提升至92.6,优于最强基线。在分层240案例编辑测试集上,MUSE实现了49.6的全目标成功率、99.9的保留率和仅0.6的非预期更改率。除了自动指标外,对比较局部编辑基线的人工评估支持与用户意图更强的对齐,下游导航代理测试表明空间稳定性更强。结合验证我们记忆设计的消融实验,这些结果确立了MUSE作为可控3D场景创作的有效框架。

英文摘要:

Text-driven 3D scene generation is a promising technique for digital content creation, embodied AI simulation, and interactive design, yet practical workflows often require refining, extending, or correcting existing scenes while preserving non-target content. Existing methods can produce realistic and structurally plausible scenes, but they generally lack editability with requirement-level state tracking, so part-level failures often lead to full-scene regeneration or manual intervention. To tackle this challenge, we formulate controllable 3D scene authoring as incremental requirement satisfaction, unifying construction and editing. In this paper, we present MUSE, a memory-grounded multi-agent framework in which an Architect compiles instructions into structured requirements, a Sculptor executes local scene operations, and an Inspector verifies each step while updating Working, Scene, and Skill Memory. To evaluate requirement-level controllability and preservation-aware editing, we introduce AuthorBench, offering 145 constrained construction cases and a 1,584-case preservation-aware editing pool paired with external structured checks. On full construction cases, MUSE improves All-Goal success from 37.9 to 80.7 and surface-constraint fulfillment from 35.0 to 92.6 over the strongest baseline. On a stratified 240-case editing test split, MUSE achieves 49.6 All-Goal success, 99.9 preservation rate, and only 0.6 unintended change rate. Beyond automated metrics, human evaluations on compared local-editing baselines support stronger alignment with user intent, and downstream navigation-proxy tests indicate stronger spatial stability. Combined with ablations validating our memory designs, these results establish MUSE as an effective framework for controllable 3D scene authoring.

↑