AI 中文总结
本文提出仅用标准VQA对训练的LUT框架,通过轨迹和步骤层面的隐式效用优化,在感知密集型视觉推理基准上性能优于现有隐式推理方法,且标注成本更低。
AI 中文摘要
多模态大语言模型已在视觉理解方面取得进展,但依赖感知的推理任务仍具挑战性。近期的隐式视觉推理方法会在回答前引入隐空间计算,但往往依赖边界框、草图或交错理由等成本较高的中间监督。这些策略关注隐状态应如何被塑造,却未明确评估隐式表示是否对最终答案有用。本文提出LUT(一种仅使用标准VQA对训练的隐式推理框架),其训练围绕两个层面的隐式效用展开:轨迹层面,提出效用感知隐式蒸馏SFT,探索与答案相关的隐式轨迹,通过信息增益筛选合格轨迹,并利用课程学习提炼更可靠、可学习的监督信号;步骤层面,提出隐式归因策略优化,在强化学习中利用答案到隐式的归因对隐式步骤进行差异化优化。在依赖感知的视觉推理基准上的实验表明,LUT的性能优于现有隐式推理方法,且在标注成本更低的情况下,与隐式-文本交错方法相比仍具竞争力。
英文摘要
Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden-space computation before answering, but they often rely on costly intermediate supervision, such as bounding boxes, sketches, or interleaved rationales. These strategies focus on how latent states should be shaped, but do not explicitly assess whether the latent is useful for the final answer. We propose LUT, a latent reasoning framework trained with only standard VQA pairs. LUT centers training on Latent Utility at two levels. At the trajectory level, we propose Utility-Aware Latent Distillation SFT, which explores answer-relevant latent trajectories, selects qualified trajectories by their information gain, and distills more reliable and learnable supervision through curriculum learning. At the step level, we propose Latent Attribution Policy Optimization, which uses answer-to-latent attribution to differentially optimize latent steps during reinforcement learning. Experiments on perception-intensive visual reasoning benchmarks show that LUT outperforms previous latent reasoning methods and remains competitive with latent-text interleaved methods with lower annotation cost.