发表机构
South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对胸部X射线视觉-语言对齐中投影视觉不匹配和语义重叠的双重歧义,提出PLRS-IC双校准框架,通过局部投影条件低秩残差和全局信息论软假阴性抑制,在九个基准上提升零样本分类、定位和分割性能。
AI 中文摘要
胸部放射学中的细粒度视觉-语言对齐能够实现零样本分类、定位和分割,而无需任务特定的标注。然而,这种对齐从根本上受到两种相互交织的歧义来源的阻碍:投影引起的视觉不匹配和与患者无关的语义重叠。首先,在局部特征层面,正位和侧位X光片对同一临床发现表现出不同的外观,使得共享的补丁-文本相似性几何结构本质上次优。这种视觉歧义在全局对比优化过程中又叠加了语义不匹配,其中实例级目标将跨患者对视为严格负样本,即使它们共享相同的阳性临床概念。为了解决这种双重歧义,我们提出了PLRS-IC,一个用于胸部X射线表示学习的统一双校准框架。在局部对齐阶段,投影条件低秩残差相似性(PLRS)通过有界、参数高效的低秩残差动态调整补丁-文本匹配以适应投影特定的流形。在全局优化阶段,信息内容校准的软假阴性抑制(IC-SFNS)利用从语料库导出的信息论先验来减轻语义重叠负样本的惩罚,而不改变原始的对比分配。在九个公开零样本基准设置上的大量实验表明,我们的框架在分类、定位和分割方面产生了一致的改进,验证了医学视觉-语言预训练中双校准的必要性。
英文摘要
Fine-grained vision-language alignment in chest radiography enables zero-shot classification, grounding, and segmentation without task-specific annotations. However, this alignment is fundamentally hindered by two intertwined sources of ambiguity: projection-induced visual mismatch and patient-agnostic semantic overlap. First, at the local feature level, frontal and lateral radiographs exhibit distinct appearances for the same clinical finding, rendering a shared patch-text similarity geometry inherently suboptimal. Compounding this visual ambiguity is a semantic mismatch during global contrastive optimization, where instance-level objectives penalize cross-patient pairs as strict negatives even when they share identical positive clinical concepts. To address this dual ambiguity, we propose PLRS-IC, a unified dual-calibration framework for chest X-ray representation learning. At the local alignment stage, Projection-Conditioned Low-Rank Residual Similarity (PLRS) dynamically adapts patch-text matching to projection-specific manifolds using a bounded, parameter-efficient low-rank residual. At the global optimization stage, Information-Content-Calibrated Soft False-Negative Suppression (IC-SFNS) leverages a corpus-derived information-theoretic prior to soften the penalty of semantically overlapping negatives without altering original contrastive assignments. Extensive experiments across nine public zero-shot benchmark settings demonstrate that our framework yields consistent improvements in classification, grounding, and segmentation, validating the necessity of dual-calibration in medical vision-language pre-training.
Comments9 pages, 4 figures, 5 tables