arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33713cs.AI

BIRD:将决策边界蒸馏为多模态大语言模型适配的推理依据

BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation

Anglin Liu, Yanlin Wu, Ruichao Chen, Yuting Zhang, Qingyuan Zeng, Pengxiang Cai, Ziqi Gong, Muchen Li, Jintai Chen

首次发表
浏览论文内容

中文总结 AI 辅助

BIRD提出一种自我改进的边界感知推理依据蒸馏框架,利用模型自身混淆定位决策边界,蒸馏有效证据以增强MLLM在专业领域的适配,在医学和图表VQA上超越现有方法。

中文摘要 AI 辅助

将通用多模态大语言模型(MLLMs)适配到专业领域需要学习领域特定的决策标准,这些标准往往依赖于看似合理答案之间的细微视觉差异。推理依据增强旨在通过额外观察或样本间比较来揭示此类证据,然而,视觉上有效的线索不一定与决策相关:它可能描述了样本之间的差异,但并未改变模型在竞争性答案之间的相对偏好。为此,我们引入了BIRD,一个自我改进的边界感知推理依据蒸馏框架,该框架利用模型特定的混淆来定位未解决的局部决策边界,并将解决这些混淆的证据蒸馏为推理依据。对于每个样本,BIRD从目标MLLM自身的表示空间中检索候选邻居,并根据其答案偏好选择最易混淆的一个。随后,它从它们的视觉差异中生成无答案偏见的候选证据,并功能性地验证哪种证据最有效地增强模型对正确答案的偏好,同时避免在样本对之间不恰当的迁移。经过验证的证据随后被蒸馏为单样本推理依据,用于标准的监督微调。在医学和图表VQA上的实验表明,BIRD在两个目标MLLM上优于竞争性的推理依据增强方法,进一步的分析则展示了混淆答案的更清晰分离以及模型匹配监督带来的更强增益。

英文摘要

Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional observations or inter-sample comparisons, yet a visually valid cue is not necessarily decision-relevant: it may describe how samples differ without changing the model's relative preference between competing answers. We therefore introduce BIRD, a self-improving Boundary-Informed Rationale Distillation framework that uses model-specific confusions to locate unresolved local decision boundaries and distills the evidence that resolves these confusions into rationales. For each sample, BIRD retrieves candidate neighbors from the target MLLM's own representation space and selects the most confusable one according to its answer preferences. It then generates answer-blind candidate evidence from their visual differences and functionally verifies which evidence most effectively strengthens the model's preference for the correct answer while avoiding inappropriate transfer across the pair. The verified evidence is then distilled into a single-sample rationale for standard supervised fine-tuning. Experiments on medical and chart VQA show that BIRD outperforms competing rationale-augmentation methods across two target MLLMs, while further analyses demonstrate clearer separation of confusable answers and stronger gains from model-matched supervision.

发表机构

  • HKUST(GZ)(香港科技大学(广州))
  • HKUST(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑