arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39363cs.CVcs.AI

重新思考多图像理解中的多图像再表示

Rethinking Multi-Image Re-Representation in Multi-Image Understanding

Gengyuan Zhang, Xiao Han, Xinyu Xie, Tong Liu, Volker Tresp

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过引入Mosaic框架和MosaicBench基准,系统比较了多图像理解中文本与视觉再表示的效果,发现视觉再表示在精确视觉证据任务中优势显著,并训练出MosaicAgent-8B智能体自主学会组合视觉操作。

中文摘要 AI 辅助

多图像理解要求多模态大语言模型(MLLMs)不仅识别单个图像的内容,还要组织分布在多个图像中的视觉证据。我们通过多图像再表示来研究这一问题,将提示式思维链推理和智能体视觉工具使用视为推理过程中重新组织视觉证据的不同方式。我们引入了Mosaic,一个通用的多图像视觉工具框架,使MLLM能够通过十种可组合的图像操作主动构建视觉中间表示。我们在现有的多图像基准测试以及MosaicBench(一个新的面向细粒度多图像理解的、以定位为核心的基准测试)上比较了五种再表示设置。我们的实验表明,文本和视觉再表示的相对优势强烈依赖于任务。视觉再表示在需要精确视觉证据的任务中尤为有效,包括假设检验、精度比较和方向敏感推理,而以高层语义内容为主的任务则显示出较小或不太一致的收益。基于这一发现,我们使用仅基于准确性和格式奖励的强化学习训练了MosaicAgent-8B以使用Mosaic。在没有演示轨迹或针对特定工具使用的奖励的情况下,该智能体学会了在多个步骤中组合视觉操作,并自发展现出多样化的问题解决模式。代码和数据将在该https网址发布。

英文摘要

Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning. We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. We compare five re-representation settings on existing multi-image benchmarks and on MosaicBench, a new grounding-focused benchmark for fine-grained multi-image understanding. Our experiments show that the relative benefits of textual and visual re-representation are strongly task-dependent. Visual re-representation is particularly effective for tasks requiring precise visual evidence, including hypothesis testing, precision comparison, and orientation-sensitive reasoning, while tasks dominated by higher-level semantic content show smaller or less consistent gains. Building on this finding, we train MosaicAgent-8B to use Mosaic with reinforcement learning using only accuracy and format rewards. Without demonstration trajectories or rewards for specific tool-use, the agent learns to compose visual operations over multiple steps and exhibits diverse problem-solving patterns unpromptedly. Code and data will be released at https://github.com/gengyuanmax/Mosaic.

发表机构

  • LMU Munich(慕尼黑大学)
  • MCML(慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑