发表机构
Technical University of Darmstadt(达姆施塔特工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ReImaGin,利用图像生成模型作为多模态LLM的灵活视觉推理机制,在六项视觉推理任务中优于文本推理和专用工具,性能提升高达25%。
AI 中文摘要
思维链推理通过使大型语言模型(LLMs)在回答问题前将问题分解为中间步骤,彻底改变了自然语言处理领域。然而,将推理局限于文本领域,对于需要直接操作视觉表示的任务存在局限性。最近的研究尝试通过外部视觉专家工具(如深度估计或目标检测模块)来增强多模态LLMs,但这些方法从根本上受限于其依赖狭窄、僵化的操作,无法灵活生成或转换视觉内容。我们提出了ReImaGin,该方法利用图像生成模型作为多模态LLMs的灵活视觉推理机制:与固定功能工具不同,图像生成模型接受自然语言指令,并能执行开放式视觉操作,如移除遮挡或从房间的多个不连续视角生成平面图。在包括多视角空间推理和碰撞预测在内的六项多样化视觉推理任务中,ReImaGin consistently优于纯文本推理和专用视觉工具基线,性能提升高达25%,展示了灵活、生成式视觉推理的优势。
英文摘要
Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to decompose problems into intermediate steps before answering. Yet confining reasoning to the textual domain presents limitations for tasks requiring direct manipulation of visual representations. Recent efforts augment multimodal LLMs with external visual expert tools such as depth estimation or object detection modules, but these remain fundamentally limited by their reliance on narrow, rigid operations that cannot flexibly generate or transform visual content. We propose ReImaGin, which leverages image generation models as a flexible visual reasoning mechanism for multimodal LLMs: unlike fixed-function tools, they accept natural language commands and can perform open-ended visual operations, like removing an occlusion or generating a floorplan from multiple disjoint views of a room. Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.
CommentsAccepted to COLM 2026. Code https://github.com/multimodal-ai-lab/reimagin and website https://hector.gr/reimagin/