超越排行榜:可信多模态视觉问答的设计经验教训
Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA
浏览论文内容
中文总结 AI 辅助
以MediaEval Medico 2025为案例,分析九个多模态问答系统,发现预训练主干的参数高效适应虽有挑战性能,但答案提升未转化为完整临床推理,结构化推理等方法更可靠,结果促使多方面改进,支持可信多模态医疗保健人工智能。
中文摘要 AI 辅助
医疗保健多模态人工智能必须结合视觉和文本证据,同时保持可靠性和可解释性。我们以MediaEval Medico 2025作为回顾性胃肠内镜检查案例研究,分析了九个有记录的问答和解释质量系统的设计选择。预训练主干的参数高效适应提供了强大的挑战性能,但答案层面的提升并不能始终转化为忠实和完整的临床推理。实施结构化推理和显式基础的方法在不同类型问题上表现出更可靠的行为,尽管证据是相关的而非基于消融的。这些结果促使进行超越词汇重叠的评估、标准化的证据关联解释、防泄漏的数据治理以及轻量级的鲁棒性和校准检查。研究结果支持基于数据融合、可解释性和弹性评估的可信多模态医疗保健人工智能。
英文摘要
Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design choices across nine documented systems for question answering and explanation quality. Parameter-efficient adaptation of pretrained backbones provides strong challenge performance, but answer-level gains do not consistently translate into faithful and complete clinical reasoning. Methods enforcing structured reasoning and explicit grounding show more reliable behavior across heterogeneous question types, although the evidence is correlational rather than ablation-based. These results motivate evaluation beyond lexical overlap, standardized evidence-linked explanations, leakage-aware data governance, and lightweight robustness and calibration checks. The findings support trustworthy multimodal healthcare AI based on data fusion, explainability, and resilient evaluation.