AI 中文总结
研究3D资产创建中用户指令模糊与工具要求精确的意图不对称问题,提出自进化智能体CLARE,通过解耦生成管道、模拟多轮交互自我进化澄清策略,在新基准测试中取得领先性能,证明主动澄清对3D执行的关键作用。
AI 中文摘要
现代3D资产创建存在基本的意图不对称问题:先进的3D工具链需要精确、可执行的参数,而普通用户通常提供模糊、未明确说明的指令。当前3D智能体将这种模糊性视为噪声,在单轮假设下盲目执行。为解决此限制,我们引入了CLARE,一种具有澄清意识和进化能力的3D智能体,它将意图不对称视为战略对话的机会。通过将生成管道解耦为四个专门的认知角色,CLARE在调用计算成本高的3D工具之前拦截并解决未明确说明的指令,以无缝执行五个不同领域的任务。关键是,CLARE通过模拟多轮交互自我进化其澄清策略。通过优化多轮奖励,智能体内化了交互效率和任务完成之间的微妙平衡。我们构建了3D-Clarify基准进行严格测试,CLARE取得了领先性能,单步和多步任务的成功率分别为60.40%和43.34%,远超现有基线。定量和定性结果都表明主动澄清是强大3D执行中缺失的关键。
英文摘要
A fundamental intent asymmetry plagues modern 3D asset creation: while state-of-the-art 3D toolchains demand precise, executable parameters, ordinary users typically provide vague, underspecified instructions. Current 3D agents treat this ambiguity as noise, defaulting to blind execution under a single-turn assumption. To address this limitation, we introduce CLARE, a clarification-aware and evolutionary 3D agent that treats intent asymmetry not as an execution error, but as an opportunity for strategic dialogue. By decoupling the generation pipeline into four specialized cognitive roles, CLARE intercepts and resolves underspecified instructions before invoking computationally expensive 3D tools to seamlessly execute tasks across five diverse domains: text-to-3D generation, single-view reconstruction, multi-view reconstruction, point cloud editing, and post-processing. Crucially, rather than relying on rigid manual rules, CLARE self-evolves its clarification policy via simulated multi-turn interactions. By optimizing a Multi-turn Reward, the agent internalizes the delicate balance between interaction efficiency and task completion. To rigorously test this, we construct 3D-Clarify, a comprehensive benchmark comprising 620 interaction scenarios with systematically injected ambiguity, missing information, and mistaken details. CLARE achieves state-of-the-art performance, with 60.40% and 43.34% success rates on single-step and multi-step tasks, respectively, more than doubling existing baselines. Both quantitative and qualitative results demonstrate that proactive clarification is the missing key to robust 3D execution. Code is available at https://github.com/xyzhu1225/CLARE.
CommentsAccepted to ACM Multimedia 2026 (ACM MM 2026). 26 pages including appendix, 8 figures. Code: https://github.com/xyzhu1225/CLARE