arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SD-MAR:通过合成数据和强化学习进行多图像分析推理

SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning

Shiyu Yuan, Sourav Sanjukta Bhabesh, Zhe Wang, Dmitriy Bespalov, Wesley Rose, Huzefa Rangwala

arXiv 2607.14333首次发表:更新:

发表机构

Stevens Institute of Technology; AGI Foundations for AWS; Amazon Web Services (AWS)(史蒂文斯理工学院; AWS的通用人工智能基金会; 亚马逊网络服务公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视觉语言模型在多图像分析推理任务中的局限,提出SD-MAR框架,通过合成数据和强化学习训练评估模型。该框架构建场景生成任务,采用GRPO-lite与BDA训练。实验表明能提升域内准确率,保持或提高域外泛化能力,还改进了模型逻辑连贯性和解释质量。

AI 中文摘要

视觉语言模型(VLMs)在感知能力方面表现出色,但在跨多个视觉状态的分析推理任务中仍存在局限性,如多图像比较、变化检测和多步视觉推理。现有基准测试很少同时要求明确的视觉比较和分析推理,这一能力未得到充分探索。为此,我们引入了SD-MAR框架来训练和评估VLMs在多图像分析推理上的表现。该框架通过可控扰动构建配对视觉场景,并生成涵盖语义变化归因和定量比较的推理任务。我们还使用GRPO-lite和反向折扣分配(BDA)对VLMs进行强化学习训练,去除KL正则化以鼓励更强的策略优化,并在形成分析结论的后期推理步骤中给予更多奖励。实验表明,在SD-MAR上对Qwen2.5-VL-7B和InternVL3-8B进行GRPO-lite微调可将域内准确率提高多达36.95%,Qwen2.5-VL-7B在SD-MAR基准测试中优于GPT-4.1。此外,域外泛化能力得以保持或提高,在多个数据集上表现良好,LLM-as-judge评估也显示出两个模型在逻辑连贯性和解释质量上的持续改进。

英文摘要

Vision Language Models (VLMs) demonstrate strong perceptual abilities but remain limited in tasks requiring analytical reasoning across multiple visual states, such as multi-image comparison, change detection, and multi-step visual inference. These capabilities are critical for real-world multimodal applications where reasoning must be grounded in systematic differences between visual contexts. However, existing benchmarks rarely require both explicit visual comparison and analytical reasoning, leaving this capability underexplored. To address this gap, we introduce SD-MAR (Synthetic Data for Multi-image Analytical Reasoning), a framework for training and evaluating VLMs on multi-image analytical reasoning. SD-MAR constructs paired visual scenarios through controlled perturbations and generates reasoning tasks spanning semantic change attribution and quantitative comparison. We further train VLMs using GRPO-lite with Backward Discounted Allocation (BDA), a reinforcement learning approach that removes KL regularization to encourage stronger policy optimization while allocating greater credit to the later reasoning steps where analytical conclusions are formed. Experiments on Qwen2.5-VL-7B and InternVL3-8B show that GRPO-lite fine-tuning on SD-MAR improves in-domain accuracy by up to 36.95%, with Qwen2.5-VL-7B outperforming GPT-4.1 on the SD-MAR benchmark. Importantly, out-of-domain generalization is preserved or improved: performance remains within 1% on MME, MMMU-Pro, and MathVista, while improving by up to 4% on MMBench. LLM-as-judge evaluation further demonstrates consistent improvements in logical coherence and explanation quality across both models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑