arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向CT视觉-语言学习的语义校准证据组合

Semantically Calibrated Evidence Composition for CT Vision-Language Learning

Guoliang You, Haifan Gong, Xiaomeng Chu

arXiv 2608.00239首次发表:更新:

AI 中文总结

该研究针对CT视觉-语言学习中局部证据与全局上下文结合的问题,提出SCOPE框架,通过语义校准实现证据组合,在CT-RATE和RadChestCT数据集上的宏AUC优于现有SOTA。

AI 中文摘要

从CT-报告对中学习可迁移表征需要将全 volume 上下文与特定解剖结构证据相结合。现有方法通常要么强调全局CT-报告对齐,要么强调细粒度解剖结构级对应。全局对齐保留了广泛的研究上下文,但未明确局部证据的贡献;而解剖结构级对齐明确锚定了局部发现,但未说明如何将独立表征的证据进行交互、获取研究级意义并贡献于全局CT表征。为解决这一差距,我们提出SCOPE(Semantic Calibration Of comPosed Evidence,组合证据的语义校准),这是一种用于CT视觉-语言学习中语义校准证据组合的框架。在特定器官的报告监督下,具有固定解剖身份的掩码引导查询从共享的未裁剪体积特征中提取上下文感知的器官证据,而无约束的全局查询则保留对全体积上下文的访问。全局查询随后驱动局部-全局耦合,将器官证据组合为统一的证据表征。组合后的证据随后使用诊断摘要进行校准,提供超出局部器官描述的研究级语义监督,并最终作为受控残差集成到与完整报告对齐的、保留上下文的全体积表征中。这一渐进式路径将局部证据与研究级语义相连接,而不会将CT表征简化为预定义的器官集合。在CT-RATE和RadChestCT上,SCOPE分别达到85.0和72.2的宏AUC,分别比之前的SOTA高出7.2和4.2,同时在线性探测和跨模态检索中也取得了显著提升。这些结果证明了语义校准证据组合的有效性。

英文摘要

Learning transferable representations from CT-report pairs requires combining whole-volume context with anatomy-specific evidence. Existing methods typically emphasize either global CT-report alignment or fine-grained anatomy-level correspondence. Global alignment preserves broad study context but leaves the contribution of localized evidence implicit, whereas anatomy-level alignment explicitly grounds local findings but does not specify how independently represented evidence should interact, acquire study-level meaning, and contribute to a global CT representation. To address this gap, we propose SCOPE (Semantic Calibration Of comPosed Evidence), a framework for semantically calibrated evidence composition in CT vision-language learning. Under organ-specific report supervision, mask-guided queries with fixed anatomical identities extract context-aware organ evidence from shared, uncropped volumetric features, while an unrestricted global query retains access to whole-volume context. The global query then drives Local-Global Coupling to compose the organ evidence into a unified evidence representation. The composed evidence is subsequently calibrated using the diagnostic summary, providing study-level semantic supervision beyond local organ descriptions, and is finally integrated as a controlled residual into a context-preserving whole-volume representation aligned with the complete report. This progressive pathway connects localized evidence with study-level semantics without reducing the CT representation to a predefined set of organs. On CT-RATE and RadChestCT, SCOPE achieves macro AUCs of 85.0 and 72.2, respectively, outperforming the previous SOTA by 7.2 and 4.2, while also yielding substantial gains in linear probing and cross-modal retrieval. These results demonstrate the effectiveness of semantically calibrated evidence composition.

Comments9 pages, 3 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑