发表机构
Tencent Hunyuan; THU; NJU; HKUST; UIUC; PKU(腾讯混元; 清华大学; 南京大学; 香港科技大学; 伊利诺伊大学厄巴纳-香槟分校; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型智能体在3D场景行动能力,提出SceneActBench基准测试,涵盖五个3D任务,通过特定指标评估智能体输出,分析不同配置得分及失败情况,为评估智能体在多对象3D场景行动提供依据。
AI 中文摘要
视觉语言模型(VLM)智能体越来越多地使用工具对3D场景采取行动,而不仅仅是描述它们。现有的3D基准测试对文本响应或单对象操作进行评分,未评估智能体在完整多对象3D场景上的行动。我们提出了SceneActBench,这是一个在统一智能体 - 环境循环下针对五个3D任务的视觉条件行动基准测试。给定PNG图像或采样视频帧以及适用时提供的3D资产,智能体在3D环境中行动。我们使用特定任务的几何指标根据隐藏的地面真值评估每个最终输出。SceneActBench由210个源实例构建的五个任务组成,产生520个任务案例,包括配对输入条件。每个任务通过一个固定的智能体循环运行以保持比较公平。在十一种专有VLM配置中,总体得分在38.6 - 50.2之间,且没有一个在所有任务中都表现良好。我们进一步分析了失败出现的位置和方式。
英文摘要
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38.6-50.2, and none performs consistently well across tasks. We further analyse where and how failures manifest.