arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Case2Flow:通过多模态检索连接患者病例与指南流程图

Case2Flow: Bridging Patient Cases and Guideline Flowcharts through Multimodal Retrieval

Jiale Wei, Yufan Chen, Alexander Jaus, Zdravko Marinov, Julian Friedrich, Simon Reiß, Jens Kleesiek, Rainer Stiefelhagen

arXiv 2608.26414首次发表:更新:

发表机构

Karlsruhe Institute of Technology; University Hospital Essen(卡尔斯鲁厄理工学院; 埃森大学医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Case2Flow任务,构建含202份流程图的FlowAtlas语料库,提出无需训练的CRISP方法优化多模态检索,提升Recall@1达18.71个百分点,为病例匹配指南流程图提供可行方案。

AI 中文摘要

医学指南包含丰富的循证决策逻辑,但临床医生很难在单份指南中找到所需的特定决策工具,更难在涵盖多种疾病和治疗方案的多份指南中定位。尽管指南段落已支持端到端问答,但流程图在决策支持中仍未得到充分利用,尽管其能编码可操作的临床路径。因此,我们提出Case2Flow任务,旨在从指南文档集合中为给定患者病例检索最相关的指南流程图。为支撑该任务,我们构建了FlowAtlas,这是一个包含从2080份医学指南中提取的202份流程图的精选语料库,以及一个合成1911对对齐病例-流程图对的流水线。我们对多模态检索方法的评估揭示了系统性失效模式,包括过度依赖关键词,以及流程图中无信息背景区域导致的虚假标记-图像块匹配。受此启发,我们提出CRISP,一种无需训练的评分方法,通过抑制无信息图像块、弱化模糊标记匹配并纳入双向查询-图像对齐来优化晚期交互检索。CRISP将Recall@1提升了高达18.71个百分点,且对已发表病例叙述的盲法医生评估提供了超出合成查询的初步可行性证据。

英文摘要

Medical guidelines encode rich, evidence-based decision logic, yet the specific decision artifact a clinician needs is hard to locate within a guideline, let alone across guidelines covering plausible diseases and treatments. While guideline passages have supported end-to-end question answering, flowcharts remain largely underused in decision support despite their ability to encode actionable clinical pathways. We therefore introduce Case2Flow, a task designed to retrieve the most relevant guideline flowchart for a given patient case from a collection of guideline documents. To support it, we construct FlowAtlas, a curated corpus of 202 flowcharts extracted from 2,080 medical guidelines, together with a pipeline that synthesises 1,911 aligned case-flowchart pairs. Our evaluation of multimodal retrieval methods reveals systematic failure modes, including overreliance on keywords and spurious token-patch matches induced by uninformative background regions in flowcharts. Motivated by this, we propose CRISP, a training-free scoring method that sharpens late-interaction retrieval by suppressing uninformative patches, discounting ambiguous token matches, and incorporating bidirectional query-image alignment. CRISP improves Recall@1 by up to 18.71 percentage points, while a blinded physician assessment on published case narratives provides preliminary feasibility evidence beyond synthetic queries.

CommentsAccepted by EMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑