文本中的去偏:相信你的视觉:面向视觉反常识推理的文本锚定跨模态迁移
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
浏览论文内容
中文总结 AI 辅助
该研究针对多模态大语言模型视觉反常识推理中语言先验干扰问题,提出文本锚定数据构建流程及后训练框架 TACT,实现无需视觉数据的跨模态去偏,提升模型视觉推理性能。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)的视觉推理能力对下游应用至关重要,尤其是反常识推理,该任务要求模型超越常见假设进行推理。近期研究主要通过增强视觉输入来改进视觉反常识推理,其假设失败源于视觉 grounding 不足。然而,我们的实证分析显示瓶颈并非视觉感知:MLLMs 已捕获相关视觉证据,正确答案存在于其解码空间中,相反,共享语言解码器在解决先验-证据冲突时更倾向于主导性语言先验,尤其针对低频事实场景。受此启发,我们首先提出文本锚定数据构建流程,其核心组件为事实频率蒸馏(FFD),用于估计常识事实的先验强度,并将验证后的反常识场景提炼为高质量文本语料。基于该语料,我们引入文本锚定后训练框架 TACT,无需任何视觉训练数据即可对共享语言解码器去偏。TACT 将遵循证据和受先验驱动的推理轨迹路由至不同优化阶段,使解码器能解决先验-证据冲突。在多个反常识视觉基准上,TACT 在保持通用能力的同时显著提升了视觉推理性能,展现了有效的文本到视觉跨模态迁移效果。
英文摘要
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models to reason beyond common assumptions. Recent studies mainly improve visual counter-commonsense reasoning by enhancing visual inputs, following the assumption that failures originate from insufficient visual grounding. However, our empirical analysis reveals that the bottleneck is not visual perception. MLLMs already capture the relevant visual evidence, and the correct answer exists in their decoding space. Instead, the shared language decoder resolves prior--evidence conflicts by favoring dominant language priors, especially for low-frequency factual scenarios. Motivated by this, we first propose a text-anchored data construction pipeline, whose core component, Fact-Frequency Distillation (FFD), estimates the prior strength of commonsense facts and distills verified counter-commonsense scenarios into a high-quality text corpus. Building upon this corpus, we introduce TACT, a text-anchored post-training framework that debiases the shared language decoder without requiring any visual training data. TACT routes evidence-following and prior-driven reasoning trajectories into different optimization stages, enabling the decoder to resolve prior--evidence conflicts. Across counter-commonsense visual benchmarks, TACT substantially improves visual reasoning while preserving general capabilities, demonstrating effective text-to-vision cross-modal transfer.
发表机构
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。