arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VABench:通过视觉演示、主动感知和度量控制衡量具身空间智能

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

Zhongbo Zhang, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, Huchuan Lu

arXiv 2609.19554首次发表:更新:

发表机构

Dalian University of Technology; Nanyang Technological University(大连理工大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VA-Bench通过视觉演示、主动感知和度量控制,评估通用多模态大模型在具身空间智能中的观察-推理-行动-修正能力,发现主动相机控制显著提升成功率,但长时程任务仍具挑战。

AI 中文摘要

空间智能不仅仅需要描述物体位置。在不完整观测下,模型必须识别并获取缺失的证据,在共同的空间框架中解释这些证据,并据此采取行动。我们引入VA-Bench来评估完整的观察-推理-行动-修正循环。通用多模态大语言模型(MLLMs)从仅RGB的演示中学习程序性上下文,主动选择相机视角,发出度量笛卡尔命令,并根据执行反馈进行修正。模型不获得特权物体姿态、预言轨迹或学习到的动作头。一个固定的、与模型无关的控制器仅执行模型指定的目标。VA-Bench包含14个基础任务族(11个单臂和3个双臂)、7个留出的几何/布局变体,以及一个长时程五物体组合轨道。我们在每个基础任务的相同20个物理验证种子上,对12个主要模型条件进行三次独立运行评估,报告终端成功率、九项轨迹级行为诊断和子任务进展。首先,表现最好的模型在标注运行中的目标定位得分为100.0%,空间关系得分为78.9%。其三次运行的任务成功宏平均仅为53.93±3.17%。其次,主动相机控制显著优于被动多视角观测的任务成功率。在一项匹配比较中,成功率从27.86%提升至57.50%。第三,留出的几何迁移可使任务成功率降低超过30个百分点。尽管有大量部分进展,但没有模型完成严格的长时程回合。因此,VA-Bench测试通用MLLMs能否将视觉演示和主动获取的证据转化为成功的具身行动。

英文摘要

Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑