发表机构
University of Strasbourg; CLCC Institut Strauss; Kiel University(斯特拉斯堡大学; CLCC施特劳斯研究所; 基尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本工作提出选项语义检索与任务特定LoRA适配,在MedReason 2026医学VQA中实现94%多项选择准确率,显著超越基线,并揭示检索与门控的边际影响。
AI 中文摘要
我们描述了对MedReason 2026挑战赛的提交,涵盖了在完全离线、容器化推理条件下的多项选择(MCQ)和开放式(OE)医学视觉问答(VQA)。我们的第一个发现是,MCQ检索必须比较答案的语义而非答案标签:标签是每个问题独立分配的,因此复制检索到的邻居的标签不会传递任何有用信息,而将当前每个选项的文本与类似训练案例中的正确答案文本进行评分,可将仅检索的准确率从20.0%提高到57.5%(在200个案例的排除检索的开发保留集上)。我们的第二个发现归因于所提交系统的准确率:保持任务特定的MCQ低秩适配(LoRA)适配器固定,改变提示内检索示例的数量k,准确率最多变化一个案例——在k=0和适配器训练时的k=1时均为187/200(93.5%),在打包运行时的默认k=3时为188/200(94.0%)——并且所提交的置信度门控覆盖在k=3的基础上没有增加净准确率,在198/200个案例中选择VLM。在最终MCQ适配器固定的情况下,检索最多改变一个案例的准确率,门控没有提供净增益。在20个开放式案例中,token-F1和RaTEScore(引用zhao2024ratescore)随着k的增加而降低,但token-F1差异的配对符号检验不显著(p≥0.29);一个单一标注者的比较发现,最终配置有6/20个错误锚点,而早期配置(在路由、适配器和提示上共同不同)有14/20个错误锚点。该系统在开发保留集上达到94.0%的MCQ准确率,在组织者的官方预评估中达到93.20%,而现成的参考基线为29.43%,同时组织者的两个开放式分数均低于该基线(地面真实一致性1.245对1.588,视觉准确率1.995对2.696,每项满分4分)。
英文摘要
We describe our submission to the MedReason 2026 challenge, covering multiple-choice (MCQ) and open-ended (OE) medical visual question answering (VQA) under fully offline, containerized inference. Our first finding is that MCQ retrieval must compare answer \emph{semantics} rather than answer labels: labels are independently assigned per question, so copying a retrieved neighbor's label transfers no useful information, whereas scoring each current option's text against correct-answer text from similar training cases raises retrieval-only accuracy from 20.0\% to 57.5\% on a 200-case retrieval-excluded development holdout. Our second finding attributes the submitted system's accuracy: holding the task-specific MCQ Low-Rank Adaptation (LoRA) adapter fixed and varying the number \(k\) of in-prompt retrieved examples changes accuracy by at most one case --- 187/200 (93.5\%) at both \(k=0\) and the adapter's training-time \(k=1\), 188/200 (94.0\%) at the packaged runtime's default \(k=3\) --- and the submitted confidence-gated override adds no net accuracy on top of \(k=3\), selecting the VLM in 198/200 cases. With the final MCQ adapter fixed, retrieval changes accuracy by at most one case, and gating provides no net gain. On 20 OE cases, token-F1 and RaTEScore~\cite{zhao2024ratescore} decrease as \(k\) grows, but paired sign tests on token-F1 differences are nonsignificant (\(p \ge 0.29\)); a single-annotator comparison found 6/20 wrong-anchor errors for the final configuration and 14/20 for an earlier configuration that jointly differed in routing, adapter, and prompting. The system reaches 94.0\% MCQ accuracy on the development holdout and 93.20\% on the organizer's official pre-evaluation, versus 29.43\% for the off-the-shelf reference baseline, while both of the organizer's open-ended scores are lower than that baseline's (ground-truth agreement 1.245 versus 1.588, visual accuracy 1.995 versus 2.696, each out of 4).
Journal refMedReason @MICCAI 2026 - Benchmarking Medical MLLM Reasoning under Domain Shift, Sep 2026, Strasbourg, France