模型指导绘图:用于统一多模态推理的自适应视觉门控
Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning
浏览论文内容
中文总结 AI 辅助
针对统一多模态模型在视觉推理中存在的问题,提出基于生成意图和视觉保真度两个内部信号的AdaViG方法,可动态评估视觉步骤,避免误导性视觉证据进入推理过程,提升了准确率并降低计算量和延迟。
中文摘要 AI 辅助
具有交错推理的统一多模态模型在视觉数学推理任务中潜力巨大。然而,生成中间视觉推理步骤有时有害且会降低推理效率。为此,本文识别出两个内部信号:生成意图和视觉保真度。基于此,提出无训练的自适应视觉门控方法AdaViG,在早期视觉生成阶段动态评估每个触发的视觉步骤,当信号弱时中止。实验表明,AdaViG提高了准确率,降低了视觉生成的计算量和时间延迟。
英文摘要
Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potential for visual mathematical reasoning tasks. However, we identify a key insight in this paradigm: generating intermediate visual reasoning steps is not always beneficial and can even be harmful, as self-generated visual steps may introduce erroneous visual evidence that misleads subsequent reasoning. Moreover, frequently triggering visual steps during reasoning incurs substantial computational and memory overhead, degrading inference efficiency. To address these accuracy and efficiency challenges, we observe that the model's internal signals can indicate whether a visual step will benefit reasoning before the entire visual generation is completed. Specifically, this work identifies two internal signals: 1) Generation Intent, which reflects whether the model has a concrete textual plan for what to draw, and 2) Visual Fidelity, which measures whether the visual generation remains grounded in the original input image. Leveraging these internal signals, we propose AdaViG, a training-free adaptive visual gating method for unified multimodal reasoning. AdaViG dynamically evaluates each triggered visual step at an early visual generation stage and aborts it when both signals are weak, thereby preventing misleading visual evidence from entering the reasoning trace while avoiding unnecessary computation. Comprehensive experiments demonstrate that AdaViG improves accuracy by up to 5.7% while reducing visual generation FLOPs by 25.0%-91.0% and wall-clock latency by 15.4%-45.6%.
发表机构
- Imperial College London(伦敦帝国理工学院)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。