arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SLVR:基于类人推理流程的结构化潜在视觉推理

SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

Albert Gao, Bing Xue, Andrea Zanette

arXiv 2610.10563首次发表:更新:

发表机构

Carnegie Mellon University; Meta(卡内基梅隆大学; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SLVR是基于Qwen2.5-VL-7B的训练框架,通过将多模态推理组织为规划等类型化潜在阶段,在多模态推理基准上取得提升,可提升细粒度视觉推理且无文本思维链的解码开销。

AI 中文摘要

多模态大语言模型(MLLMs)在回答视觉推理问题时,往往依赖语言先验而非与任务相关的视觉证据。文本思维链推理可通过鼓励模型将视觉问题分解为中间证据搜寻步骤,部分缓解该问题,但自回归生成这些步骤会增加推理成本。潜在推理避免显式理由生成,但现有方法对中间状态的编码控制有限,难以对规划、定位和证据选择施加独立监督。我们提出结构化潜在视觉推理(SLVR),这是一种训练框架,通过将多模态推理组织为规划、定位、证据选择和推理整合的类型化潜在阶段,弥合显式思维链与潜在推理之间的差距。SLVR首先通过掩盖揭示答案的文本,并将正确答案与视觉上合理的干扰项进行对比,训练模型依赖图像;随后将推理组织为规划、定位、证据选择和整合的潜在阶段,用对应的信号(规划、边界框、视觉证据和最终理由)监督每个阶段。这使潜在推理具备显式功能结构,同时在推理时避免生成文本链。SLVR基于Qwen2.5-VL-7B构建,在多模态推理基准上持续提升,在MMVP上绝对提升9.4,在BLINK Relation上绝对提升14.2,同时在V*、MathVista和ChartQA上也有改进。这些结果表明,结构化潜在监督可提升细粒度视觉推理,且无文本思维链的解码开销。项目页面可在此处访问。

英文摘要

Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, but generating these steps autoregressively increases inference cost. Latent reasoning avoids explicit rationale generation, but existing approaches provide limited control over what intermediate states encode, making it difficult to impose separate supervision for planning, grounding, and evidence selection. We propose Structured Latent Visual Reasoning (SLVR), a training framework that bridges explicit chain-of-thought and latent reasoning by organizing multimodal reasoning into typed latent stages for planning, grounding, evidence selection, and reasoning integration. SLVR first trains the model to rely on the image by masking answer-revealing text and contrasting the correct answer with visually plausible distractors. It then organizes reasoning into latent stages for planning, grounding, evidence selection, and integration, supervising each stage with the corresponding signal: plans, boxes, visual evidence, and final rationales. This gives latent reasoning an explicit functional structure while avoiding generated textual chains at inference time. Built on Qwen2.5-VL-7B, SLVR improves consistently across multimodal reasoning benchmarks, with absolute gains of +9.4 on MMVP and +14.2 on BLINK Relation, as well as improvements on V*, MathVista, and ChartQA. These results suggest that structured latent supervision can improve fine-grained visual reasoning without the decoding overhead of textual CoT. Project page is available \href{https://bogao-code.github.io/SLVR/}{here}.

CommentsAccepted by NeurIPS 2026.Project page \href{https://bogao-code.github.io/SLVR/}

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑