arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MindTopo:基础模型能否在拓扑空间中进行推理?

MindTopo: Can Foundation Models Reason in Topological Space?

Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Jianwen Lyu, Zihan Wang, Reuben Tan, Jianfeng Gao, Ruohan Zhang, Yining Hong, Jiajun Wu, Manling Li

arXiv 2609.11900首次发表:更新:

发表机构

Northwestern University; Microsoft Research; Stanford University(西北大学; 微软研究院; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MindTopo基准,从认知科学和形式拓扑出发,评估基础模型在连续性等五个拓扑属性上的推理与规划能力,发现推理优于规划且均远低于人类水平。

AI 中文摘要

空间推理不仅依赖于距离、角度和形状等度量属性,还依赖于在连续变形下保持不变的拓扑关系。认知科学将这些关系视为空间理解的基础,然而基础模型的评估主要集中于度量或视角依赖的关系。我们提出了MindTopo,一个基于认知科学和形式拓扑学中五个属性(连续性、分离性、顺序性、封闭性和纽结性)的拓扑直觉基准。MindTopo在两个认知层面评估每个属性。推理要求模型识别拓扑关系或推断它们如何变化。规划则将基础模型实例化为一个闭环智能体,其策略选择环境动作。MindTopo包含11,030个实例,涵盖13种程序化生成的任务类型,难度可控。我们对14个多模态大语言模型(MLLMs)进行了基准测试,并研究了增强图像和视频生成的智能体配置,包括在规划设置中的3个视频生成模型。每个MLLM在推理上的表现均优于规划,且最佳模型仍远低于观察到的人类表现。在Qwen3-VL-2B-Instruct上,监督微调和强化学习对推理的提升大于规划。生成的观测保留了局部线索并达到合理的终点,但审计的轨迹并未可靠地遵循环境动态或在转换中保持拓扑结构。我们的网站位于此https URL。

英文摘要

Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metric or viewpoint-dependent relations. We introduce MindTopo, a benchmark of topological intuition across five properties grounded in cognitive science and formal topology: continuity, separation, order, enclosure, and knots. MindTopo evaluates each property at two cognitive levels. Reasoning asks a model to identify topological relations or infer how they change. Planning instantiates a foundation model as a closed-loop agent whose policy selects environment actions. MindTopo contains 11,030 instances across 13 procedurally generated task types with controllable difficulty. We benchmark 14 MLLMs and study agent configurations augmented with image and video generation, including 3 video generative models in planning settings. Every MLLM performs better on reasoning than on planning, and the best-performing model remains far below observed human performance. On Qwen3-VL-2B-Instruct, supervised fine-tuning and reinforcement learning improve reasoning more than planning. Generated observations retain local cues and reach plausible endpoints, but audited rollouts do not reliably follow environment dynamics or preserve topology across transitions. Our website is at https://mind-topo.github.io/

CommentsPreprint version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑