迈向基于人类先验的物理感知空间智能:一项自动驾驶试点研究
Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- Johns Hopkins University(约翰霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出空间智能网格(SIG)架构及SIGBench基准,通过结构化网格编码场景几何与物理先验,在多模态大模型少样本学习中显著提升视觉-空间智能评估的稳定性与全面性,并支持自动驾驶场景的类人VSI任务。
AI中文摘要:
如何在基础模型中整合并验证空间智能仍是一个开放的挑战。当前实践通常用纯文本提示和VQA式评分来代理视觉-空间智能(VSI),这掩盖了几何信息,引入了语言捷径,并削弱了对真正空间技能的归因。我们引入了空间智能网格(SIG):一种结构化的、基于网格的架构,显式编码对象布局、对象间关系以及物理基础先验。作为文本的补充通道,SIG为基础模型推理提供了场景结构的忠实、组合式表示。基于SIG,我们推导出SIG信息化的评估指标,用于量化模型的内在VSI,将空间能力与语言先验分离。在与最先进的多模态LLM(如GPT和Gemini系列模型)的少样本上下文学习中,与仅VQA表示相比,SIG在所有VSI指标上产生了一致更大、更稳定、更全面的增益,表明其作为学习VSI的数据标注和训练架构的潜力。我们还发布了SIGBench,一个包含1.4K驾驶帧的基准,标注了真实SIG标签和人类注视轨迹,支持自动驾驶场景中基于网格的机器VSI任务和注意力驱动的类人VSI任务。
英文摘要:
How to integrate and verify spatial intelligence in foundation models remains an open challenge. Current practice often proxies Visual-Spatial Intelligence (VSI) with purely textual prompts and VQA-style scoring, which obscures geometry, invites linguistic shortcuts, and weakens attribution to genuinely spatial skills. We introduce Spatial Intelligence Grid (SIG): a structured, grid-based schema that explicitly encodes object layouts, inter-object relations, and physically grounded priors. As a complementary channel to text, SIG provides a faithful, compositional representation of scene structure for foundation-model reasoning. Building on SIG, we derive SIG-informed evaluation metrics that quantify a model's intrinsic VSI, which separates spatial capability from language priors. In few-shot in-context learning with state-of-the-art multimodal LLMs (e.g. GPT- and Gemini-family models), SIG yields consistently larger, more stable, and more comprehensive gains across all VSI metrics compared to VQA-only representations, indicating its promise as a data-labeling and training schema for learning VSI. We also release SIGBench, a benchmark of 1.4K driving frames annotated with ground-truth SIG labels and human gaze traces, supporting both grid-based machine VSI tasks and attention-driven, human-like VSI tasks in autonomous-driving scenarios.