发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究联合设计以智能体为中心的领域特定语言aDSL与角色专业化多智能体系统,通过“规划-执行-评审”循环提升3D内容创建的鲁棒性等,在文本到形状等任务上优于LLM基线。
AI 中文摘要
程序化表示为3D内容创建提供了极具吸引力的范式,支持细粒度编辑、可解释性和显式结构控制。然而,依赖大型语言模型(LLM)编写3D程序的智能体工作流程较为脆弱,常无法将高级意图转化为一致的低级几何结构。我们将这种脆弱性归因于现有程序化接口与LLM推理优势之间的不匹配——LLM更擅长语义结构和空间关系的推理,而非易出错的数值选择。本文中,我们联合设计了一种以智能体为中心的领域特定语言(aDSL)和角色专业化的多智能体系统,以缩小这一差距。aDSL通过强调可组合性和空间推理,衔接语义逻辑与几何约束;它使智能体能够通过关系运算符而非易出错的绝对坐标来操作几何结构。基于aDSL构建的无训练多智能体系统遵循“规划-执行-评审”循环,用于分解请求、合成代码,并利用执行反馈迭代修复错误和约束违反。实验表明,这种联合设计提升了鲁棒性、可控性和对用户意图的忠实度。我们的方法在文本到形状、图像到形状任务上优于现有基于LLM的基线,同时保留了显式结构、可编辑性和可解释性;还支持下游应用,如关节物体创建和结构化场景合成。我们的代码可在该https URL获取。
英文摘要
Programmatic representations provide a compelling paradigm for 3D content creation, enabling fine-grained edits, interpretability, and explicit structural control. Yet, agentic workflows that rely on large language models (LLMs) to author 3D programs remain brittle, often failing to translate high-level intent into consistent low-level geometry. We attribute this fragility to a mismatch between existing programmatic interfaces and the reasoning strengths of LLMs, which favor semantic structure and spatial relations over fragile numeric choices. In this paper, we jointly design an Agent-centric Domain-Specific Language (aDSL) and a role-specialized multi-agent system to close this gap. aDSL bridges semantic logic and geometric constraints by emphasizing composability and spatial reasoning; it enables agents to manipulate geometry through relational operators instead of brittle absolute coordinates. Building on aDSL, our training-free multi-agent system follows a Plan-Execute-Critic loop to decompose requests, synthesize code, and iteratively repair errors and constraint violations using execution feedback. Experiments show that this co-design improves robustness, controllability, and faithfulness to user intent. Our method outperforms prior LLM-based baselines on text-to-shape and image-to-shape tasks while preserving explicit structure, editability, and interpretability. It also enables downstream applications such as articulated object creation and structured scene composition. Our code is available at https://github.com/sig-pku/aDSL.