发表机构
Mohamed bin Zayed University of Artificial Intelligence; Khalifa University; Indian Institute of Technology Delhi(穆罕默德·本·扎耶德人工智能大学; 哈里发大学; 印度理工学院德里分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究医学视觉语言模型测试时模态泛化问题,提出MoBE框架,通过熵引导动态路由与专家级贝叶斯适应相结合,无需参数更新,在多基准测试中比现有方法平均准确率有显著提升,实现无训练专家适应以增强模态泛化。
AI 中文摘要
医学视觉语言模型(MVLMs)有望实现广泛的零样本泛化,但其在面对未见模态和领域时可靠性会崩溃,而这正是临床稳健性最为关键之处。为填补这一差距,我们从专家混合(MoE)角度重新审视测试时模态泛化,并提出问题:专家在推理时能否无需任何优化就进行路由和适应?我们识别出测试时的一个基本专业化 - 泛化困境,盲目聚合模态专家会稀释特定模态知识,而选择一个高度自信的专家在分布变化时存在不匹配风险。为解决此问题,我们提出MoBE:一个完全无需优化的框架,在测试时执行动态专家选择和适应。MoBE将MoE设置中的熵引导动态路由与专家级贝叶斯适应相结合,使专家无需梯度更新就能在线更新置信度并适应。无需参数更新,MoBE通过测试时路由和在线统计增强静态MVLM,在可见、未见和异构医学基准上比现有TTA方法平均准确率分别提高+4.72、+7.17和+4.3,突出了无训练专家适应对稳健模态泛化的有效性。
英文摘要
Medical vision-language models (MVLMs) promise broad zero-shot generalization, yet their reliability collapses when confronted with unseen modalities and domains, precisely where clinical robustness matters most. To address this gap, we revisit test-time modality generalization from the perspective of Mixture-of-Experts (MoE) and ask: can experts route-and-adapt without any optimization during inference? We identify a fundamental specialization-generalization dilemma at test time, where blindly aggregating modality experts dilutes modality-specific knowledge, while selecting one highly confident expert risks mismatch under shift. To address this, we propose MoBE: a fully optimization-free framework that performs dynamic expert selection and adaptation at test time. MoBE combines entropy-guided dynamic routing in MoE settings with expert-wise Bayesian adaptation, enabling experts to update their confidence and adapt online without gradient updates. Without parametric updates, MoBE augments a static MVLM with test-time routing and online statistics, achieving average accuracy gains of +4.72, +7.17, and +4.3 over state-of-the-art TTA methods across seen, unseen, and heterogeneous medical benchmarks, highlighting the effectiveness of training-free expert adaptation for robust modality generalization.
CommentsAccepted to MICCAI 2026