发表机构
HKUST(GZ); Tencent; AI Thrust; TEG AIPD(香港科技大学(广州); 腾讯; AI thrust; TEG AIPD)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出VibeWorlding框架,构建VWE-BENCH基准与VibeWorlding-Gym框架,发现当前前沿MLLMs在3D开放世界构建任务中表现不佳,经RL训练的VibeWorlder模型可提升性能,旗舰模型VibeWorlder-30B-A3B表现最优。
AI 中文摘要
根据用户查询构建交互式3D开放世界具有重要意义,但现有方法主要基于理想化的简单查询进行评估,难以系统分析和比较多模态智能体如何理解用户意图、使用3D工具、对文本与视觉3D世界信息进行推理。为此,我们提出VibeWorlding,这是一个用于基准测试和训练氛围世界构建(vibe worlding)智能体的统一框架,该智能体是一种多模态智能体,可在多轮智能体-环境交互过程中自主推断用户意图、规划场景布局、调用3D工具、对多模态反馈进行反思。为实现这一目标,我们首先构建VWE-BENCH,该基准包含2616个高质量3D资产、323个人工标注的种子3D世界、6828个反向合成的多模态用户查询,分为带有真值的验证查询和带有精心设计评分标准的未验证查询。此外,我们开发VibeWorlding-Gym,这是一个联合多模态强化学习(RL)后训练框架,整合了(1)将资产检索、编辑和图像渲染统一为MCP工具的沙箱环境,(2)结合物理可行性与意图满足验证的基于评分标准的验证器,支持公平的模型评估和可扩展的多模态RL奖励服务。我们的实验表明,当前前沿多模态大语言模型(MLLMs)远未解决氛围世界构建智能体任务,即使GPT-5.5和Qwen3.8-Max的成功率也低于60%,并将瓶颈追溯到精确的3D世界编辑。我们进一步发现,RL训练可缓解这一弱点,使开源MLLMs甚至能超越闭源前沿模型:我们的VibeWorlder-8B可与前沿MLLMs相媲美,而我们的旗舰模型VibeWorlder-30B-A3B在所有评估模型中达到最佳的整体Pass@1指标。
英文摘要
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.
Commentspreprint