arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多模态大语言模型的符合规则的视觉空间规划

Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

Yu Chen, Ting Lei, Yaoyi Li, Jia Cai, Zhecen Wu, Yang Liu

arXiv 2608.20237首次发表:更新:

发表机构

Wangxuan Institute of Computer Technology, Peking University; Yinwang Intelligent Technology Co., Ltd(北京大学王选计算机研究所; 银湾智能科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多模态大语言模型的规则遵循空间规划问题,构建了RuleMaze基准,提出语言-逻辑-函数混合方法和解耦多模态规划(DMP),提升了规则遵循度与规划成功率。

AI 中文摘要

多模态大语言模型(MLLMs)将语言推理与视觉感知相结合,但它们在显式或未见规则约束下执行视觉空间规划的能力仍未得到充分探索。该场景要求模型共同理解空间布局、解释自然语言规则并据此规划有效动作。为解决这一空白,我们引入RuleMaze,这是一个可控基准,要求MLLMs在遵守不同复杂度的自然语言规则的同时在迷宫中导航。RuleMaze通过要求准确感知、规则解释和受限动作规划来隔离符合规则的空间规划。为实现可扩展且系统的规则构建,我们提出语言-逻辑-函数混合方法,该方法自动生成自然语言规则并将其转换为逻辑表示和可执行验证器,消除了手动规则工程。为改进规则遵循和泛化,我们引入解耦多模态规划(DMP),它通过可解释的推理原语将感知、执行和规则验证分离开来。通过解耦这些组件,DMP促进了对更复杂和未见规则的系统泛化,同时提供透明的中间规划轨迹。实验表明,与端到端文本规划基线相比,DMP显著提高了规则遵循度和规划成功率。总体而言,RuleMaze为研究MLLMs中基于接地和可解释的基于规则的空间规划建立了一个原则性基准。代码可在this https URL获取。

英文摘要

Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored. This setting requires models to jointly understand spatial layouts, interpret natural-language rules, and plan valid actions accordingly. To address this gap, we introduce RuleMaze, a controllable benchmark in which MLLMs must navigate mazes while obeying natural-language rules of varying complexity. RuleMaze isolates rule-compliant spatial planning by requiring accurate perception, rule interpretation, and constrained action planning. To enable scalable and systematic rule construction, we propose Language-Logic-Function Hybridization, which automatically generates natural-language rules and translates them into logical representations and executable validators, eliminating manual rule engineering. To improve rule following and generalization, we introduce Disentangled Multimodal Planning (DMP), which separates perception, execution, and rule verification through interpretable reasoning primitives. By disentangling these components, DMP facilitates systematic generalization to more complex and previously unseen rules, while providing transparent intermediate planning traces. Experiments demonstrate that DMP substantially improves rule compliance and planning success compared to end-to-end textual planning baselines. Overall, RuleMaze establishes a principled benchmark for studying grounded and interpretable rule-based spatial planning in MLLMs. Code is available at https://github.com/oceanflowlab/RuleMaze.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑