超越表面模仿:多模态上下文学习中推理路径对齐的对比建模
Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning
- Sun Yat-Sen University(中山大学)
- Peking University(北京大学)
- Pengcheng Laboratory(鹏城实验室)
- Tianjin University of Science and Technology(天津科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态上下文学习中表面模仿导致推理路径不对齐的问题,提出对比演示建模与响应条件检索及对齐控制器相结合的框架,显著提升MLLM性能,尤其VQA。
AI中文摘要:
上下文学习(ICL)被广泛应用于多模态大语言模型(MLLMs)中,并在广泛的多模态任务上取得了强劲的性能。然而,现有的多模态ICL方法往往依赖于对上下文演示的表面级模仿,这使得MLLMs难以将其响应与给定多模态输入所需的推理路径对齐。这一局限在复杂多模态任务中更为突出,从而限制了MLLM性能的进一步提升。为解决此问题,我们提出了一种新的多模态ICL框架,该框架将对比演示建模与MLLMs的自我精炼能力相结合。具体而言,我们的框架通过在同一输入下显式对比次优响应与更优响应,并附带揭示响应应如何被精炼的推理路径,来重新构建每个演示。这种对比性表述使通向期望响应的推理路径更加明确,并引导MLLM超越表面模仿。此外,由于有效的精炼依赖于当前响应,我们引入了一种响应条件检索机制,以选择推理路径与当前响应更相关的演示。另外,我们使用一个轻量级对齐控制器来预测响应质量,并决定是否需要进一步精炼。在三种类型的多模态任务上的实验表明,所提出的框架持续提升了MLLM的性能,尤其在视觉问答(VQA)上取得了显著提升。
英文摘要:
In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface level imitation of in-context demonstrations, making it difficult for MLLMs to align their responses with the reasoning path required by the given multimodal input. This limitation becomes more pronounced in complex multimodal tasks, thereby restricting further improvements in MLLM performance. To address this issue, we propose a new multimodal ICL framework that combines contrastive demonstration modeling with the self-refinement capability of MLLMs. Specifically, our framework reformulates each demonstration by explicitly contrasting a suboptimal response with a better response under the same input, together with a reasoning path that reveals how the response should be refined. This contrastive formulation makes the reasoning path toward the desired response more explicit and guides the MLLM beyond superficial imitation. Furthermore, because effective refinement depends on the current response, we introduce a response-conditioned retrieval mechanism to select demonstrations whose reasoning paths are more relevant to the current response. In addition, we use a lightweight alignment controller to predict response quality and determine whether further refinement is needed. Experiments on three types of multimodal tasks show that the proposed framework consistently improves MLLM performance, with particularly notable gains on visual question answering (VQA).