AlloEgo-VLM:在视觉语言模型中消除 allocentric( allocentric 即 allocentric 参考框架,又称 allocentric 坐标系,指以环境为中心的参考框架)与 egocentric( egocentric 即 egocentric 参考框架,又称自我中心坐标系,指以观察者自身为中心的参考框架)参考框架的歧义
AlloEgo-VLM: Disambiguating Allocentric and Egocentric Reference Frames in Vision-Language Models
- National Yang Ming Chiao Tung University(国立阳明交通大学)
- Institute of Computer Science and Engineering(工程与计算机科学学院)
- College of Artificial Intelligence(人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对VLMs理解空间语义时 allocentric 与 egocentric 参考框架的歧义问题,构建数据集AlloEgo-View并开发框架AlloEgo-VLM,经NVIDIA Isaac Sim平台验证其在具身机器人开放式物体搜索任务中的有效性。
AI中文摘要:
本研究探讨了视觉语言模型(Vision-Language Models,VLMs)在理解空间语义时面临的歧义挑战。空间认知受认知心理学、空间科学和文化背景的影响,常为物体赋予方向性;然而,自然语言对空间关系的描述常省略明确的参考框架,导致语义歧义,可能给具身智能机器人带来严重错误。现有VLMs因参考框架和物体方向的训练不足,常产生不一致的响应。为解决该问题,我们构建了新数据集AlloEgo-View,包含(图像、查询、特定视角答案)三元组,捕获 allocentric 和 egocentric 视角下的关键物体关系;特定视角描述遵循结构化空间表示,标注详细场景描述、参考与目标物体、其方向、参考框架及视角类型。基于AlloEgo-View,我们开发框架AlloEgo-VLM,用于消除 allocentric 和 egocentric 参考框架的歧义,即使在歧义查询下也能工作,且可通过监督微调轻松集成到现有VLMs中。此外,我们将该框架部署到NVIDIA Isaac Sim 内的具身机器人平台,验证其在开放式物体搜索任务中的现实可行性。实验凸显了当前VLMs在处理特定视角查询时的局限性,并证明了AlloEgo-VLM强大的歧义消除能力。
英文摘要:
This study investigates the challenge of ambiguity faced by Vision-Language Models (VLMs) in understanding spatial semantics. Spatial cognition, shaped by cognitive psychology, spatial science, and cultural context, often assigns directionality to objects. However, natural language descriptions of spatial relations frequently omit explicit reference frames, leading to semantic ambiguity and potentially serious errors for embodied AI robots. Existing VLMs, due to insufficient training on reference frames and object orientations, often produce inconsistent responses. To address this issue, we construct a new dataset, AlloEgo-View, comprising (image, query, view-specific answer) triplets that capture key object relations from both allocentric and egocentric perspectives. The view-specific descriptions follow a structured spatial representation that annotate detailed scene descriptions, reference and target objects, their orientations, reference frames, and view types. Building on AlloEgo-View, we develop AlloEgo-VLM, a framework to disambiguate allocentric and egocentric reference frames, even under ambiguous queries, and to be easily integrated into existing VLMs via supervised fine-tuning. Furthermore, we deploy our framework onto an embodied robotic platform within NVIDIA Isaac Sim to validate its real-world feasibility in open-ended object searching tasks. Experiments highlight the limitations of current VLMs in handling view-specific queries and demonstrate the strong disambiguation ability of AlloEgo-VLM.