优势引导门:重塑基于视觉的空间智能的开放式推理
Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence
浏览论文内容
中文总结 AI 辅助
本文针对MLLMs开放式推理易出错的问题,提出优势引导门框架,结合蒙特卡洛价值评估与两类优势门,构建Reasoning-Tree-160k数据集并经两阶段学习,有效提升了基准MLLMs的视觉空间推理性能。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在复杂空间场景理解与推理任务中展现出巨大潜力,但其开放式推理过程易出现决策错误与误差累积,导致答案质量不稳定。为解决该问题,本文提出优势引导门框架,可在推理过程中动态干预并修正偏差。具体而言,本文将逐步推理建模为有限时域决策过程,并在推理树上引入蒙特卡洛价值评估以提供中间监督信号;该框架包含Step-Advantage Gate(步骤优势门)与Trajectory-Advantage Gate(轨迹优势门),分别用于动态选择高价值推理步骤与高质量完整推理轨迹。训练阶段,本文利用多分支采样生成的推理树对门进行监督学习,结合共享参数初始化与任务特定头以实现跨任务的鲁棒性与多样性;推理阶段,模型贪婪选择高价值前缀推理步骤,并根据问题类型选择最优推理头,从而显著提升最终答案的准确性。此外,本文构建了Reasoning-Tree-160k数据集并在其上进行两阶段学习,大量实验表明该优势引导门框架可有效提升基准MLLMs在基于视觉的空间理解与推理任务中的性能,代码已公开供研究使用。
英文摘要
Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning process is prone to decision errors and error accumulation, leading to instability in answer quality. To address this, we propose an advantage-guided gating framework that dynamically intervenes in and corrects deviations during the reasoning process. Specifically, we model step-by-step reasoning as a finite-horizon decision process and introduce Monte Carlo value evaluation on the reasoning tree to provide intermediate supervision signals. The framework includes Step-Advantage Gate and Trajectory-Advantage Gate, which dynamically select high-value reasoning steps and high-quality complete reasoning trajectories, respectively. During training, we perform supervised learning for the gates using reasoning trees generated via multi-branch sampling, and combine shared-parameter initialization with task-specific heads to achieve cross-task robustness and diversity. During inference, the model greedily selects high-value prefix reasoning steps while choosing the optimal reasoning head based on the problem type, thereby significantly improving the accuracy of the final answer. Furthermore, we constructed the Reasoning-Tree-160k dataset and performed two-stage learning on it. Extensive experiments demonstrate that this advantage-guided gating framework effectively enhances the performance of benchmark MLLMs in visual-based spatial understanding and reasoning tasks. The code is open to the public for research: https://github.com/LingLin-ll/Advantage-Guided-Gate.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Suzhou Institute for Advanced Research, USTC(中国科学技术大学苏州高等研究院)
- Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高级智能与计算研究所)
- East China Normal University(华东师范大学)
- National University of Singapore(新加坡国立大学)
- Durham University(杜伦大学)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。