arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于混合格式医学视觉问答的带排列稳定专家的候选扩展路由

Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA

Hai-Dang Nguyen, Huy-Hieu Pham

arXiv 2609.00959首次发表:更新:

发表机构

College of Engineering and Computer Science, VinUniversity; VinUni-Illinois Smart Health Center, VinUniversity; Center for Innovations in Health Sciences, VinUniversity(VinUniversity工程与计算机科学学院; VinUniversity VinUni-Illinois智能健康中心; VinUniversity健康科学创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对混合格式医学VQA的两类失效问题,提出含排列稳定视觉-语言专家与稀疏候选扩展路由的方法,在1403例数据上提升准确率,开放式输出全部符合模式,验证了候选扩展的路由增益与输出契约有效性。

AI 中文摘要

混合格式医学视觉问答(VQA)需要稳定的选项选择和机器可读的自由文本输出。两种格式的失效方式不同:多项选择预测会随选项符号或位置变化,而临床合理的开放式回答若序列化格式错误则无法通过自动化评估。我们通过答案文本记忆模块、排列稳定的视觉-语言专家(permutation-stabilized vision--language expert)以及稀疏候选扩展路由(sparse candidate-expanding router)解决这两个挑战。循环调度遵循现有研究;我们的贡献是将专家前2(expert top-2)与记忆模块、专家前1(expert top-1)一同设为可路由候选。在含1403个案例的回顾性内部分析中,该扩展使匹配的二元路由准确率从88.95%提升至91.73%(提升2.78个百分点;95%置信区间1.57--3.99),同时挽救56个错误、出现17个退化情况。最优覆盖度(Oracle coverage)从90.31%升至96.15%,最终提交的配置在同一回顾性划分上达到92.23%。对于开放式问题,严格生成与确定性防护机制产出475/475个符合模式的面向参与者输出,无需修复、重试或硬门控失效。可视化消融实验显示存在显著文本依赖性。候选扩展提供主要的受控路由增益;开放路径证据确立了输出契约有效性,而非医学使用或部署中的临床正确性。

英文摘要

Mixed-format medical visual question answering (VQA) requires stable option selection and machine-readable free-text output. The two formats fail differently: multiple-choice predictions can change with option symbols or positions, while clinically plausible open answers can fail automated evaluation when serialization is malformed. We address both challenges with an answer-text memory, a permutation-stabilized vision--language expert, and a sparse candidate- expanding router. The cyclic schedule follows prior work; our contribution is to make expert top-2 a routable candidate alongside memory and expert top-1. On a 1,403-case retrospective internal analysis, this expansion improves a matched binary router from 88.95% to 91.73% (+2.78 percentage points; 95% CI 1.57--3.99), with 56 rescued errors and 17 regressions. Oracle coverage rises from 90.31% to 96.15%, and the final submitted configuration reaches 92.23% on the same retrospective split. For open questions, strict generation and deterministic guards produce 475/475 schema- valid participant-facing outputs without repair, retry, or hard-gate failure. Visual ablations reveal substantial textual dependence. Candidate expansion supplies the principal controlled routing gain; open-path evidence establishes output-contract validity rather than clinical correctness in medical use or deployment.

Comments10 pages, 3 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑