发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出以视觉为中心的无训练多智能体系统ViSculpt,通过Blender GUI直接编辑现有三维网格,可遵循指令执行局部编辑并保留资产整体特征,为语言驱动的三维编辑提供了新范式。
AI 中文摘要
三维几何编辑是图形流程中关键却耗时的环节,需要创作者在复杂专业软件中将创作意图转化为精确操作。大语言模型(LLM)在基于脚本的三维创作中已展现潜力,但脚本生成不太适用于对任意现有网格的感知驱动编辑,此类编辑需保持视觉接地性并保留未改动区域。我们提出一种以视觉为中心、无需训练的多智能体系统,该系统通过模拟人类艺术家的迭代工作流程,直接在Blender中编辑现有三维网格。该系统不生成脚本或重新生成几何,而是通过Blender图形用户界面(GUI)运行:多模态LLM智能体观察视口、推理当前网格状态,并通过模拟用户交互执行局部编辑。在精心筛选的基准上开展的实验初步证明,这种智能体方法能够遵循自然语言指令、执行代表性的局部网格编辑,且能保留输入资产的整体特征。我们的结果凸显了语言驱动三维编辑的互补范式:在原生三维编辑工作流中直接对现有网格进行原位修改,将此项工作视为专业图形软件中以视觉为中心的智能体几何编辑的探索性一步。
英文摘要
3D geometry editing is a critical yet labor-intensive part of the graphics pipeline, requiring artists to translate creative intent into precise operations in complex professional software. Large language models (LLMs) have shown promise for script-based 3D creation, but script generation is less suited to perception-driven editing of arbitrary existing meshes, where execution must remain visually grounded and untouched regions should be preserved. We present a \emph{visual-centric}, training-free multi-agent system that edits existing 3D meshes directly in Blender by emulating the iterative workflow of human artists. Rather than generating scripts or regenerating geometry, our system operates through the Blender GUI: multimodal LLM agents observe the viewport, reason about the current mesh state, and execute localized edits through simulated user interactions. Experiments on a curated benchmark provide initial evidence that this agentic approach can follow natural language instructions, perform representative localized mesh edits, and preserve the overall identity of the input asset. Our results highlight a complementary regime for language-driven 3D editing: direct in-place modification of existing meshes within the native 3D editing workflow. We view this work as an exploratory step toward visual-centric agentic geometry editing in professional graphics software.