arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言推理向量能否增强多模态推理能力?

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang

arXiv 2609.31140首次发表:更新:

发表机构

Southeast University; University of California, Los Angeles(东南大学; 加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LIFT方法,通过从基础LLM提取推理向量并注入VLM语言侧,无需重训练即可恢复多模态模型退化推理能力,实验证明LLM来源向量更有效。

AI 中文摘要

大多数视觉-语言模型(VLMs)是通过在预训练的大型语言模型(LLMs)上扩展视觉模块和多模态对齐而构建的。然而,这种多模态扩展往往会削弱基础LLM中原本编码的语言侧推理能力。虽然基础LLM在扩展后仍保留可用的推理能力,但对齐后的VLM本身无法可靠地访问这种能力。因此,恢复VLM中退化的推理能力,从基础LLM寻求帮助比单独依赖VLM更为有效。基于此,我们提出LIFT(语言侧推理促进与迁移),一种轻量级向量干预方法,无需重新训练骨干网络即可将推理能力从基础LLM迁移至VLM。LIFT将推理向量定义为带有显式推理轨迹的推理者路径与不带推理轨迹的求解者路径之间答案标记的隐藏状态差异,并将这些向量注入目标VLM的语言侧激活中。LIFT还支持可学习的向量自适应,同时保持VLM骨干网络冻结。我们在两个VLM上、六个推理基准中评估LIFT,在匹配协议下比较从基础LLM和从对齐VLM提取的推理向量。结果表明,从LLM提取的向量始终优于从VLM提取的向量,证实基础LLM是恢复推理能力的更有效来源。LIFT通过轻量级语言侧干预部分恢复了退化的推理能力。进一步分析表明,推理向量影响中间推理行为,而非仅仅改变最终答案。源代码即将发布。

英文摘要

Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering the degraded reasoning capability in VLMs would benefit more from seeking help from the base LLM than from the VLM alone. Motivated by this, we propose LIFT (Language-side reasonIng Facilitation and Transfer), a lightweight vector-intervention method that transfers reasoning capability from the base LLM to the VLM without retraining the backbone. LIFT defines Reasoning Vectors as answer-token hidden-state differences between a Reasoner path with an explicit reasoning trace and a Solver path without it, and injects these vectors into language-side activations of the target VLM. LIFT further supports learnable vector adaptation while keeping the VLM backbone frozen. We evaluate LIFT on two VLMs across six reasoning benchmarks, comparing Reasoning Vectors extracted from the base LLM and from the aligned VLM under matched protocols. Results show that LLM-derived vectors consistently outperform VLM-derived vectors, confirming that the base LLM is a more effective source for recovering reasoning. LIFT partially recovers degraded reasoning through lightweight language-side interventions. Further analyses show that Reasoning Vectors influence intermediate reasoning behavior rather than merely altering final answers. The source code will be released soon.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑