arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自我在空间中:无人机具身智能中的自我意识与空间认知基准测试

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan, Yujie Li, Mugen Peng, Wenjia Xu

arXiv 2607.12477首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; Hunan Normal University(北京邮电大学; 湖南师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对无人机具身智能中自我意识与空间认知问题,引入SIS - Bench基准,通过多维度层次评估,揭示现有模型局限,探索运动感知表征提升性能,强调自我意识重要性,提供新基准和实证依据。

AI 中文摘要

自主无人机系统越来越依赖多模态大语言模型在复杂现实环境中运行。此类具行情景不仅需要理解周围空间,还需保持智能体自身的连贯表征。但现有无人机方法和基准主要以环境为中心,智能体自我意识隐含。为此引入SIS - Bench基准,沿空间和自我两个维度及感知、记忆、推理三级层次组织评估。它含4856个问答对,源于1646个真实无人机视频的13项任务。评估发现当前模型有局限,空间认知与自我意识失衡且性能下降。进而探索运动感知表征,实验表明其能提升性能,强调自我意识对具身空间智能的重要性,提供新基准和实证证据。

英文摘要

Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification. Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels. Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks. Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.

CommentsWebsite:https://choucisan.github.io/publications/self-in-space ; Code:https://github.com/IntelliSensing/Self-in-Space

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑