保持符号投影布局:用于视觉-语言模型中以物体为中心的空间推理的符号投影布局
Keep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language Models
- Kyung Hee University(庆熙大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SymPL通过将以物体为中心的空间推理转化为符号布局形式,提升视觉-语言模型在多视角场景下的推理性能和鲁棒性。
AI中文摘要:
具有视角意识的空间推理涉及从特定视角理解空间关系——无论是以自我为中心(观察者为中心)还是以物体为中心(物体为中心)。尽管视觉-语言模型(VLMs)在以自我为中心的环境中表现良好,但当从以物体为中心的视角进行推理时,其性能会下降,因为必须从场景中物体的视角推断空间关系。在本研究中,我们通过引入符号投影布局(SymPL)框架来解决这一未被充分探索的挑战,该框架将以物体为中心的推理重新表述为VLMs能够很好地处理的符号布局形式。通过利用四个关键因素——投影、抽象、二分法和定位,SymPL将以物体为中心的问题转换为结构化的符号布局表示。广泛的实验表明,这种重新表述在以物体为中心和以自我为中心的任务中都显著提高了性能,增强了在视觉错觉和多视角场景下的鲁棒性,并且每个组件都对这些增益做出了关键贡献。这些结果表明,SymPL为解决复杂的视角意识空间推理提供了有效且原理性的方法。
英文摘要:
Perspective-aware spatial reasoning involves understanding spatial relationships from specific viewpoints-either egocentric (observer-centered) or allocentric (object-centered). While vision-language models (VLMs) perform well in egocentric settings, their performance deteriorates when reasoning from allocentric viewpoints, where spatial relations must be inferred from the perspective of objects within the scene. In this study, we address this underexplored challenge by introducing Symbolic Projective Layout (SymPL), a framework that reformulates allocentric reasoning into symbolic-layout forms that VLMs inherently handle well. By leveraging four key factors-projection, abstraction, bipartition, and localization-SymPL converts allocentric questions into structured symbolic-layout representations. Extensive experiments demonstrate that this reformulation substantially improves performance in both allocentric and egocentric tasks, enhances robustness under visual illusions and multi-view scenarios, and that each component contributes critically to these gains. These results show that SymPL provides an effective and principled approach for addressing complex perspective-aware spatial reasoning.