arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MetaSpace:面向具身智能体空间认知的变形测试

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

Gengyang Xu, Dongwei Xiao, Yiteng Peng, Shuai Wang

arXiv 2608.07533首次发表:更新:

发表机构

Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MetaSpace是评估具身智能体空间认知的框架,基于变形测试原则生成测试用例,在三个场景中检测到SOTA MLLM驱动智能体的90422个空间认知错误,其SC分数远低于人类基准。

AI 中文摘要

具身智能体是通过物理身体与环境交互的智能实体。当前对具身智能体的评估主要依赖两种范式:(1)人工标注的视觉问答(VQA)对;(2)高级任务完成指标,如导航或操作的成功度。前者耗费人力且受标注质量差异影响,后者可能掩盖关键漏洞,使智能体通过次优方式完成任务或违反安全规范,从而隐藏安全风险与低效性。鉴于空间认知是执行具身任务的基石,迫切需要评估具身智能体在任务执行中是否具备稳健的空间认知能力。受软件工程中变形测试原则启发,我们提出MetaSpace,这一用于评估智能体空间认知的新型框架。通过利用真实执行轨迹衍生的时空多模态状态,MetaSpace基于逻辑规则和物理定律定义的预定义变形关系(MRs)自动生成测试用例。关键在于,我们将这些MRs编码为逻辑编程语言(Prolog)中的可执行规则。违反这些关系即表明空间认知存在故障。我们在三个具身场景中的实证评估显示,MetaSpace成功检测到最先进(SOTA)多模态大语言模型(MLLM)驱动的智能体中的90422个空间认知错误。我们引入空间认知(SC)分数来量化性能,结果表明所有SOTA智能体的平均分数在0.44至0.52之间,显著低于人类基准的0.96。

英文摘要

An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents primarily relies on two paradigms: (1) manually annotated Visual Question Answering (VQA) pairs and (2) high-level task completion metrics, such as success in navigation or manipulation. The former is labor-intensive and subject to variability in annotation quality. The latter may obscure critical vulnerabilities, allowing agents to complete tasks through suboptimal means or safety violations, thereby concealing safety risks and inefficiencies. Given that spatial cognition is the cornerstone for executing embodied tasks, there is a pressing need to assess whether embodied agents possess robust spatial cognition during task execution. Inspired by metamorphic testing principles in software engineering, we propose MetaSpace, a novel framework designed to evaluate the spatial cognition of agents. By leveraging spatiotemporal multimodal states derived from real execution trajectories, MetaSpace automatically generates test cases based on predefined metamorphic relations (MRs) grounded in logical rules and physical laws. Crucially, we encode these MRs as executable rules in a logic programming language (Prolog). Violations of these relations indicate failures in spatial cognition. Our empirical evaluation across three embodied scenarios demonstrates that MetaSpace successfully detects 90,422 spatial cognition errors in state-of-the-art (SOTA) MLLM-driven agents. We introduce the Spatial Cognition (SC) score to quantify performance. Results indicate that all SOTA agents achieve average scores between 0.44 and 0.52, significantly lower than the human benchmark of 0.96.

Comments30 pages, 17 figures. Published in Proceedings of the ACM on Programming Languages (OOPSLA1)

Journal refProceedings of the ACM on Programming Languages, 10, OOPSLA1 (April 2026), 343-372

DOI:10.1145/3798212

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑