发表机构
School of Computer Science and Technology, Beijing Institute of Technology; Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology; School of Computer Science and Engineering, Beihang University(北京理工大学计算机科学与技术学院; 哈尔滨工业大学社会计算与交互机器人研究中心; 北京航空航天大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有具身探索范式泛化能力不足的问题,构建了基准VGEBench及逻辑驱动状态机框架,实验发现现有VLMs在语义知识转物理执行和长程状态跟踪上存在挑战。
AI 中文摘要
视觉语言模型(VLMs)在静态视觉识别和高级语义推理方面已展现出令人印象深刻的能力,但当前的具身探索范式仍严重依赖人类标注轨迹的模仿学习,这极大限制了智能体的泛化能力。实现通用自主具身智能体的关键瓶颈在于可泛化视觉接地探索:即无需手册或特定训练,通过将抽象世界知识主动接地到细粒度视觉可供性,就能操作新设备的能力。然而现有基准无法评估该能力,它们通常依赖显式文档和标注轨迹,忽略了功能设备操作必需的动态假设-交互-细化过程。为弥合这一差距,我们引入VGEBench,这是一个用于评估VLMs可泛化视觉接地探索能力的综合基准。与静态数据集不同,我们构建了逻辑驱动状态机(Logic-Driven State Machine)框架,该框架模拟多轮交互循环,迫使智能体通过主动视觉感知和反馈驱动的修正来达成目标。实验结果表明,现有VLMs在将语义知识转化为物理执行以及维持长程状态跟踪方面面临重大挑战。
英文摘要
Recent advancements in Vision-Language Models (VLMs) have demonstrated impressive capabilities in static visual recognition and high-level semantic reasoning. However, current embodied exploration paradigms still heavily rely on imitation learning from human-annotated trajectories, which severely limits agents' generalization ability. The key bottleneck of realizing general autonomous embodied agents lies in Generalizable Visually Grounded Exploration: the ability to operate novel devices without manuals or specific training by actively grounding abstract world knowledge into fine-grained visual affordances. Yet, existing benchmarks fail to evaluate this capability: they generally rely on explicit documents and annotated trajectories, neglecting the dynamic Hypothesis-Interaction-Refinement process essential for functional device operation. To bridge this gap, we introduce VGEBench, a comprehensive benchmark designed to evaluate the generalizable visually grounded exploration capabilities of VLMs. Unlike static datasets, we construct a Logic-Driven State Machine framework. This framework simulates multi-turn interaction loops, compelling agents to achieve goals by active visual perception and feedback-driven correction. Experimental results demonstrate that existing VLMs face significant challenges in translating semantic knowledge into physical execution and maintaining long-horizon state tracking.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026