AI 中文总结
研究针对多模态思维链推理在小模型中存在的问题,提出视觉显著性引导蒸馏方法,利用多模态大语言模型注意力图生成扰动图像,经奇异值分解提取引导向量指导层间蒸馏,实验证明该方法能改进推理生成和答案推理。
AI 中文摘要
多模态思维链(CoT)推理通过逐步推理整合视觉和文本线索。在令牌预算有限的小模型中,模态交互融合常常抑制微小的跨模态差异。特别是当不同图像与相同文本或不同文本与相同图像配对时,多模态CoT往往难以应对,融合后此类输入几乎无法区分。本研究提出视觉显著性引导蒸馏(VSSD)。VSSD利用多模态大语言模型的注意力图生成捕获任务敏感特征方向的扰动图像,然后应用奇异值分解提取主导引导向量以指导层间蒸馏。在ScienceQA和M³CoT上的实验表明,VSSD改进了推理生成和答案推理。代码可在该https网址获取。
英文摘要
Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budgets, modality-interaction fusion often suppresses tiny cross-modal differences. In particular, multimodal CoT often struggles when different images pair with identical text or different texts pair with an identical image, making such inputs nearly indistinguishable after fusion. This study proposes Visual Saliency Steering Distillation (VSSD). VSSD leverages the attention maps of multimodal large language models to generate perturbed images that capture task-sensitive feature directions, and then applies singular value decomposition to extract dominant steering vectors to guide inter-layer distillation. Experiments on ScienceQA and M$^3$CoT demonstrate that VSSD improves rationale generation and answer inference. The code is available at https://github.com/BGWH123/VSSD.
DOI:10.1109/ICASSP55912.2026.11463006