arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缓解临床分布偏移下医学视觉-语言模型的类别尾部覆盖不足问题

Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

Mushir Akhtar, M. Tanveer

arXiv 2607.28696首次发表:更新:

发表机构

Indian Institute of Technology Indore(印度理工学院印多雷分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对临床分布偏移下医学VLMs的类别尾部覆盖不足问题,提出CALCoDe方法,在多个偏移场景和骨干网络上实现了所有设置下边际和最差类别覆盖度均达0.95的最优表现。

AI 中文摘要

医学视觉-语言模型(VLMs)在临床分布偏移后可保持较高的观测边际覆盖度,但会大幅遗漏个别疾病类别,受影响的类别随采集协议和骨干网络几何结构变化,因此源域流行度无法可靠反映该失效情况。现有局部共形方法和尾部感知共形方法分别适配测试邻域和源域频率尾部,未对保留的类别级覆盖失效进行建模。本文提出Class-Tail Adaptive Localized Conformal Deferral(CALCoDe),这是一种针对冻结医学VLMs的事后可靠性层:交叉拟合的验证预测识别出存在覆盖不足风险的类别,不相交的校准集估计其类别条件尾部阈值;CALCoDe通过单侧最大值将每个受保护阈值与局部共形阈值结合,所得集合包含局部规则接纳的所有标签,额外保护仅限定于验证识别的类别;独立校准的支持度审计会对异常值支持不足的案例执行弃权(不执行)操作。在每个受保护类别内接受样本可交换的假设下,CALCoDe在预设防护水平下提供有限样本覆盖度,且包含对应的局部共形集合;本文在两个皮肤病学偏移场景(HAM10000到ISIC 2019、HAM10000到PAD-UFES-20)及四个冻结VLM骨干网络(BiomedCLIP、OpenAI CLIP ViT-B/32、PubMedCLIP ViT-B/32、MedSigLIP-448)上,对标准共形基线和近期VLM专用共形方法进行评估,结果显示CALCoDe是唯一在全部8个场景中观测边际和最差类别接受覆盖度均达到0.95的方法;在HAM10000到ISIC 2019场景中,其平均最差类别接受覆盖度为0.970,而sTACP为0.926、LCP-VLM为0.864。

英文摘要

Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual disease class. The affected class varies with acquisition protocol and backbone geometry, so source prevalence does not reliably reveal the failure. Existing localized and tail-aware conformal methods respectively adapt to test neighborhoods and source-frequency tails, leaving held-out class-wise coverage failure unmodeled. We introduce Class-Tail Adaptive Localized Conformal Deferral (CALCoDe), a post-hoc reliability layer for frozen medical VLMs. Cross-fitted validation predictions identify classes at risk of undercoverage, and a disjoint calibration split estimates their class-conditional tail thresholds. CALCoDe combines each protected threshold with a localized conformal threshold using a one-sided maximum. The resulting set contains every label admitted by the localized rule, with additional protection confined to validation-identified classes. An independently calibrated support audit defers cases with insufficient inlier support. Under exchangeability among accepted examples within each protected class, CALCoDe provides finite-sample coverage at the prespecified guard level and contains the corresponding localized conformal sets; coverage on shifted external cohorts is evaluated empirically. Among standard conformal baselines and recent VLM-specific conformal methods evaluated across two dermatology shifts (HAM10000 to ISIC 2019 and HAM10000 to PAD-UFES-20) and four frozen VLM backbones (BiomedCLIP, OpenAI CLIP ViT-B/32, PubMedCLIP ViT-B/32, and MedSigLIP-448), CALCoDe is the only approach whose observed marginal and worst-class accepted coverage both reach 0.95 in all eight settings. On HAM10000 to ISIC 2019, its average worst-class accepted coverage is 0.970, compared with 0.926 for sTACP and 0.864 for LCP-VLM.

Comments26 pages; supplementary material included

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑