WorldAuditBench:基于多模态智能体的交互式3D世界审计
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
- UC Santa Barbara(加州大学圣塔芭芭拉分校)
- MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
- MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出WorldAuditBench基准,包含213个3D异常审计任务,评估多模态智能体耦合行动与视觉推理的能力,结果显示成功率远低于人类,揭示了当前局限。
AI中文摘要:
随着交互式3D世界越来越多地被用于研究智能行为,开发高效的流水线来识别这些模拟环境中的异常(如漂浮物体、可穿越的墙壁或与周围场景不一致的物体)变得尤为重要。多模态人工智能系统,包括视觉语言模型(VLM)和视觉语言行动模型(VLA),已显示出自动化此任务的潜力。然而,3D世界审计十分复杂,需要紧密耦合两种不同的能力:行动能力,即系统地、高效地导航3D世界并搜索异常;以及视觉推理能力,即理解环境并从多模态观察中识别异常。多模态智能体能否有效耦合这两种能力,利用视觉推理识别潜在异常,同时采取行动验证这些异常,在很大程度上仍未得到探索。在本文中,我们引入了WorldAuditBench,一个用于3D世界审计的基准,包含13个基于Unreal Engine 5和此http URL构建的环境中的213个异常任务,涵盖五类异常。我们在固定的探索预算下,使用两种审计范式评估了五个前沿模型:基于VLA的探索后接基于VLM的异常识别,以及端到端的VLM智能体,其中视觉推理直接指导行动选择。在评估的模型和两种范式中,成功率范围为6.6%至42.3%,远低于人类表现(83.4%)。通过世界审计任务,WorldAuditBench为研究多模态智能体如何在交互式3D环境中耦合行动和视觉推理提供了一个测试平台,同时突显了它们在探索过程中收集和解释证据能力的当前局限性。
英文摘要:
As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.