发表机构
Mohamed bin Zayed University of Artificial Intelligence; Nanjing University(穆罕默德·本·扎耶德人工智能大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MLLMs依赖外部工具感知图表等视觉内容的问题,提出TwSG框架,通过两阶段训练实现细粒度视觉推理,降低延迟并提升准确率与鲁棒性。
AI 中文摘要
能够结合图像进行思考的多模态大语言模型(MLLMs)通常依赖外部工具实现细粒度感知,但这种依赖会引入显著的推理延迟,且无法有效解决空间结构差距——这是文本密集型、具有结构关联的视觉内容(如图表和视觉表格)中的根本挑战,这类内容中严格的相对空间排列约束着文本元素。若不借助外部工具,标准MLLMs难以完成此类细粒度视觉推理任务。为解决这些问题,我们提出了Think with Structured Grounding(TwSG),这是一种新型细粒度图像感知框架,旨在将复杂图像的工具使用能力内化到模型中。TwSG将多步推理和微裁剪的优势整合到推理阶段的单次高效前向传播中。具体而言,我们利用MLLM在真实答案的引导下识别关键区域,随后提示教师模型生成高质量的视觉问答(VQA)数据。这些基于区域的细粒度监督信号随后被蒸馏回全图像表示中。我们的训练流程包含两个阶段:(1)冷启动监督微调(SFT)阶段,使用带有聚焦区域描述的多轮数据,以培养复杂推理和错误恢复能力;(2)由新型过程奖励机制TL-GRPO驱动的强化微调(RFT)阶段,该机制鼓励策略性推理。在各类MLLM架构上开展的大量实验表明,TwSG可降低推理延迟,同时大幅提升准确率和鲁棒性,赋予模型原生的细粒度区域描述能力与灵活推理能力。
英文摘要
Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this reliance introduces significant inference latency and fails to effectively resolve the spatial-structural gap-a fundamental challenge in text-dense and structurally relational visuals (e.g., charts and visual tables) where strict relative spatial arrangements bind textual elements. Without external tools, standard MLLMs struggle with such fine-grained visual reasoning tasks. To address these issues, we propose Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model. TwSG distills the benefits of multi-step reasoning and micro-cropping into a single efficient forward pass during inference. Specifically, we use an MLLM to identify key regions guided by ground-truth answers, and then prompt a teacher model to generate high-quality visual question-answering (VQA) data. These fine-grained, region-based supervisory signals are subsequently distilled back into the full-image representation. Our training pipeline consists of two stages: (1) a cold-start supervised fine-tuning (SFT) phase using multi-turn data with focused area descriptions to foster complex reasoning and error recovery; and (2) a reinforcement fine-tuning (RFT) phase driven by a novel process reward mechanism, TL-GRPO, which encourages strategic reasoning. Extensive experiments across various MLLM architectures demonstrate that TwSG reduces inference latency while substantially improving accuracy and robustness, endowing models with native fine-grained region description and flexible reasoning capabilities.
CommentsManuscript