CheXGround:用于 grounded 纵向胸部 X 线片解读的解剖区域标记
CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation
- Ajman University(阿治曼大学)
- University of Birmingham(伯明翰大学)
- Khalifa University(哈利法大学)
- University of Western Australia(西澳大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 CheXGround,一种基于区域的纵向胸部 X 线片语言模型,通过时序区域-短语对齐预训练目标关联区域标记与临床文本,在多项放射学任务上较基线模型提升性能,验证了解剖层面组织纵向证据的有效性。
AI中文摘要:
近期,放射学多模态语言模型在胸部 X 线片报告生成、视觉问答及时序推理方面已取得显著进展。纵向胸部 X 线片解读需对比序列检查以描述变化,而视觉定位旨在将临床语言与图像局部证据建立关联。尽管纵向建模与视觉定位各自推动了放射学语言模型的发展,但局部视觉证据如何支撑纵向解读仍未得到充分探索。本文提出 CheXGround,这是一种基于区域的纵向胸部 X 线片语言模型,通过对应解剖区域表示配对研究。CheXGround 从当前及先前的 X 线片中提取解剖区域,将其编码为时序增强的感兴趣区域(ROI)标记,并在生成过程中与全局时序图像上下文结合。为将这些区域标记与临床文本关联,本文提出时序区域-短语对齐这一预训练目标,用于对齐时序解剖表示与报告局部短语。本文在单研究及纵向视觉问答(VQA)、纵向发现生成、时序 grounded VQA 及解剖定位任务上对 CheXGround 进行评估。在所有任务中,CheXGround 较近期基线模型提升了临床语言质量、时序推理及定位准确率。研究结果表明,在解剖层面组织纵向证据是 grounded 放射学语言建模的有效表示。项目页面:this https URL
英文摘要:
Recent radiology multi-modal language models have made substantial progress in chest X-ray report generation, visual question answering, and temporal reasoning. While longitudinal chest X-ray interpretation compares sequential examinations to describe change, visual grounding aims to connect clinical language with localized image evidence. Although longitudinal modeling and visual grounding have each advanced radiology language models, how localized visual evidence can support longitudinal interpretation remains under-explored. We introduce CheXGround, a region-grounded longitudinal chest X-ray language model that represents paired studies through corresponding anatomical regions. CheXGround extracts anatomical regions from current and prior radiographs, encodes them as temporally enhanced Region-of-Interest (ROI) tokens, and combines them with global temporal image context during generation. To connect these region tokens with clinical text, we propose Temporal Region--Phrase Alignment, a pretraining objective that aligns temporal anatomical representations with localized report phrases. We evaluate CheXGround on single-study and longitudinal Visual Question Answering (VQA), longitudinal findings generation, temporal grounded VQA, and anatomical grounding. Across these tasks, CheXGround improves clinical language quality, temporal reasoning, and localization accuracy over recent baselines. Our results suggest that organizing longitudinal evidence at the anatomical level is a strong representation for grounded radiology language modeling. Project page: https://adonaydem.github.io/chexground-website