发表机构
Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对MedFrameQA数据集,提出目标对齐的直接答案SFT是最强稳健适配系列,其在保留报告准确率上优于冻结基线且稳定性好,还可迁移至其他骨干网络,旨在引导研究聚焦稳健优化而非架构复杂度。
AI 中文摘要
多帧医学视觉问答(VQA)似乎需要越来越复杂的适配方法:控制器式推理、感知定位的重排序、静态难负样本混合以及分阶段延续,从基本原理来看这些方法似乎都可行。我们在MedFrameQA数据集上检验一个更简单的竞争假设:一旦在固定划分、匹配预算、重复种子和校准等条件下控制评估,那些与基准最终答案目标紧密对齐的方法应是最强的稳健适配系列。我们对比了基于控制器的方法、支架演化、静态混合监督、重延续变体以及仅直接答案的监督微调(SFT)。最强的稳健系列是基于MedGemma-1.5-4B的仅解码器直接答案SFT。经验证,该系列在保留报告准确率上较冻结基线有显著提升,且在重复种子和匹配控制下表现极为稳定,确保我们的结论反映的是真正的系列级稳健性,而非孤立的超参数峰值。此外,事后校准可有效修复置信度估计且不损害准确率,核心方法还能一致迁移到Qwen2.5-VL-3B等次级骨干网络。因此,主要结论并非复杂辅助机制胜出,而是目标对齐的直接答案SFT是我们在MedFrameQA上发现的最强稳健适配系列。通过建立这个强大且极简的基线,我们希望将研究界的关注点转向根本上稳健的优化,而非架构复杂性。
英文摘要
Multi-frame medical VQA appears to reward increasingly complex adaptation: controller-style inference, localization-aware reranking, static hard-negative mixing, and staged continuation all appear plausible from first principles. We test a simpler competing hypothesis on MedFrameQA: methods that remain tightly aligned with the benchmark's final answer objective should be the strongest \emph{robust} adaptation family once evaluation is controlled across fixed splits, matched budgets, repeated seeds, and calibration. We compare controller-based methods, scaffold evolution, static mixed supervision, continuation-heavy variants, and direct answer-only supervised fine-tuning (SFT). The strongest robust family is direct decoder-only answer SFT on MedGemma-1.5-4B. Empirically, this family yields substantial improvements in held-out report accuracy over frozen baselines while remaining remarkably stable across repeated seeds and matched controls, ensuring our claims reflect true family-level robustness rather than an isolated hyperparameter peak. Furthermore, post-hoc calibration effectively repairs confidence estimation without compromising accuracy, and the core approach transfers consistently to secondary backbones like Qwen2.5-VL-3B. The main result is therefore not that a complex auxiliary mechanism wins, but that objective-aligned direct answer SFT is the strongest robust adaptation family we found for MedFrameQA. By establishing this strong, minimalist baseline, we hope to redirect community focus toward fundamentally robust optimization rather than architectural complexity.
CommentsPresented at the CVPR 2026 Workshop on Multimodal Foundation Models for Biomedicine: Challenges and Opportunities