发表机构
School of Electrical and Information Engineering, Beihang University; Institute of Unmanned System, Beihang University; Atmanity Inc.; Megvii Research; University of California at Merced(北京航空航天大学电子信息工程学院; 北京航空航天大学无人系统研究院; Atmanity公司; 旷视研究院; 加州大学默塞德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有3D场景编辑方案的缺陷,提出Hash-Atlas网络及基于LLM的CE3D++方法,实现2D编辑与3D重建解耦,可调度30种视觉工具,支持3D/4D场景交互式文本驱动编辑。
AI 中文摘要
近期,基于视觉-语言预训练模型的图像内容操作相关研究已被有效扩展至文本驱动的3D场景编辑领域。然而,现有的3D场景编辑方案仍存在一定缺陷,阻碍其作为交互式设计工具的进一步发展:这类方案通常遵循固定的输入模式,限制了文本输入的灵活性;此外,其编辑能力受限于单个或少数2D视觉模型,且需要复杂的流水线设计以将这些模型集成到3D重建过程中。为解决上述问题,我们提出Hash-Atlas网络,该网络将3D场景编辑重新表述为对2D图集图像的操作,从而实现2D编辑与3D重建过程的解耦工作流。在此基础上,我们引入一种基于对话的3D场景编辑方法,命名为CE3D++,其以大语言模型(LLM)为核心,支持用户任意文本输入并解读其意图,进而促进对相应视觉模型的自主调用。此外,我们通过对运动对象施加运动约束,将CE3D++扩展至单目4D场景,并创建与编辑任务相关的轨迹数据集以进一步微调LLM,这使得规模较小的LLM能够准确调度多达30种不同的视觉工具。实验结果表明,CE3D++可有效集成多种视觉模型以实现多样的视觉编辑效果,具备强大的场景理解与多轮对话能力。源代码与训练模型可在该https URL获取。
英文摘要
Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for 3D scene editing still have certain shortcomings, hindering their further development as interactive design tools. Such schemes typically adhere to fixed input patterns, limiting flexibility in text input. Furthermore, their editing capabilities are constrained by a single or a few 2D visual models and require intricate pipeline design to integrate these models into 3D reconstruction processes. To address the aforementioned issues, we propose the Hash-Atlas network, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes. Building on this foundation, we introduce a dialogue-based 3D scene editing approach, termed CE3D++, which is centered on a large language model (LLM) that allows arbitrary textual input from users and interprets their intentions, subsequently facilitating the autonomous invocation of the corresponding visual models. Additionally, we extend CE3D++ to monocular 4D scenes by imposing motion constraints on moving objects and further fine-tuning the LLM by creating a trajectory dataset related to editing tasks, which enables the smaller LLM to schedule up to 30 different visual tools accurately. Experimental results demonstrate that CE3D++ effectively integrates multiple visual models to achieve diverse visual editing effects, possessing strong scene comprehension and multi-round dialog capabilities. The source codes and trained models are available at https://github.com/Fangkang515/CE3D.
CommentsProject Website: https://sk-fun.fun/CE3D