发表机构
Technion – Israel Institute of Technology; Ben-Gurion University of the Negev(以色列理工学院; 内盖夫本-古里安大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对胸部X线检索的目标不匹配问题,提出CXR-Retrieve基准与标签感知对比微调方法,在双病理组合和否定查询的Precision@5上较CXR-CLIP分别提升8.5和22.0个百分点。
AI 中文摘要
大型胸部X线摄影档案难以检索,因为大多数研究仅与自由文本报告配对,而非结构化临床注释。视觉语言模型为文本到图像检索提供了自然接口,但当前生物医学模型主要针对报告到图像匹配进行优化,而非满足简短临床搜索查询,这造成了目标不匹配:模型可能检索到与查询中词语相关的图像,却无法满足完整临床约束,尤其是对于“肺不张且无肺炎”这类连词和否定表述。我们提出CXR-Retrieve,这是一个用于组合式胸部X线文本到图像检索的结构化基准。该基准包含来自MIMIC-CXR-JPG官方测试集拆分的5159张测试图像,以及145个文本查询,涵盖单一和组合发现、阳性和阴性情况。相关性定义为检索到的图像是否满足所有断言的病理学约束,而非是否匹配配对报告。我们进一步提出用于临床检索的标签感知对比微调目标,该方法吸引具有兼容断言病理学约束的图像-文本对,包括共享的确认缺失,同时明确排斥矛盾对。从域内CXR-CLIP检查点开始,我们的方法在双病理组合任务上将Precision@5较CXR-CLIP提升了8.5个百分点,在否定查询上提升了22.0个百分点。这些结果表明,可靠的胸部X线检索需要训练目标不仅要对提及的发现进行建模,还要对其临床断言方式进行建模。
英文摘要
Large chest radiography archives are difficult to search because most studies are paired only with free-text reports rather than structured clinical annotations. Vision-language models offer a natural interface for text-to-image retrieval, but current biomedical models are primarily optimized for report-to-image matching rather than for satisfying short clinical search queries. This creates an objective mismatch: a model may retrieve images related to words in the query while failing to satisfy the full clinical constraint, especially for conjunctions and negations such as ``atelectasis and no pneumonia.'' We introduce CXR-Retrieve, a structured benchmark for compositional chest X-ray text-to-image retrieval. The benchmark contains 5,159 test images from the official test-split of MIMIC-CXR-JPG and 145 textual queries spanning single and conjunction findings, both positive and negative. Relevance is defined by whether a retrieved image satisfies all asserted pathology constraints, rather than by whether it matches a paired report. We further propose a label-aware contrastive fine-tuning objective for clinical retrieval. Our method attracts image-text pairs with compatible asserted pathology constraints, including shared confirmed absences, while explicitly repelling contradictory pairs. Starting from the in-domain CXR-CLIP checkpoint, our method improves Precision@5 over CXR-CLIP by 8.5 percentage points on two-pathology conjunctions and by 22.0 percentage points on negation queries. These results show that reliable chest X-ray retrieval requires training objectives that model not only which findings are mentioned, but also how they are clinically asserted.