arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37374cs.CVcs.MM

MG-Thinker:面向多图像推理定位的双轴自反思方法

MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding

Heyu Huang, Chi Chen, Zonghao Guo, Yuhua Li, Maosong Sun, Ruixuan Li

首次发表
浏览论文内容

中文总结 AI 辅助

MG-Thinker提出一种基于双轴自反思的后训练强化学习框架,利用25K带思维链标注的数据集,通过BiA-DAPO分解优势,实现多图像推理定位的最先进性能与泛化提升。

中文摘要 AI 辅助

强化学习(RL)近期在多模态推理中带来了显著的性能提升,为细粒度视觉感知开辟了一条有前景的路径。然而,对于多图像推理定位(MRG)任务——即在真实世界的多图像情境中进行推理以实现像素级精确定位——现有的基于RL的方法忽视了该范式固有的两个特性:由粗到细的层次化推理模式,以及异构分布的任务-样本难度。在本工作中,我们提出了MG-Thinker,一种后训练RL框架,它推进了一种具有这种层次化推理的新MRG范式,并辅以一个精心整理的包含25K条MRG数据的数据集,该数据集带有任务自适应的思维链(CoT)标注,能够在得出结论前引出多视角证据。为了解决异构的任务-样本难度,我们进一步提出了双轴DAPO(BiA-DAPO),它通过两种互补机制,将rollout优势沿组内信号轴和组间能力轴进行分解,这两种机制均基于我们定义的候选池,以获得稳定的组级统计量。大量实验表明,MG-Thinker在多图像推理定位上达到了最先进的性能,同时持续提升了在多图像理解及多样化多模态基准上的泛化能力。

英文摘要

Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pixel-precise localization, existing RL-based approaches overlook two characteristics intrinsic to this paradigm: a coarse-to-fine hierarchical reasoning pattern, and heterogeneously distributed task--sample difficulties. In this work, we present MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarchical reasoning, supported by a curated 25K MRG dataset with task-adaptive Chain-of-Thought (CoT) annotations that elicit multi-perspective evidence before conclusion. To remedy the heterogeneous task--sample difficulties, we further propose Bi-Axial DAPO (BiA-DAPO), which decomposes rollout advantages along an intra-group signal axis and an inter-group competence axis through two complementary mechanisms, both grounded on our defined candidate pool for stable group-level statistics. Extensive experiments show that MG-Thinker achieves state-of-the-art performance on multi-image reasoning grounding while consistently improving generalization across multi-image understanding and diverse multimodal benchmarks.

发表机构

  • Huazhong University of Science and Technology(华中科技大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑