测试时自适应何时能帮助零样本CT视觉语言模型?
When Can Test-Time Adaptation Help Zero-Shot CT Vision-Language Models?
浏览论文内容
中文总结 AI 辅助
研究零样本3D CT视觉语言模型中测试时自适应(TTA)的作用,通过控制诊断分析表明其具有条件性,引入CARVE方法,在基础模型有判别力时能带来一致改进,确立多标签TTA问题及CARVE为基数感知解决方案。
中文摘要 AI 辅助
3D CT视觉语言模型(VLMs)以零样本方式根据文本提示对异常进行分类,有助于跨机构部署。然而,真实CT扫描通常包含多个同时出现的异常,零样本多标签预测在分布变化下的可靠性尚不清楚。现有测试时自适应(TTA)方法不适用于零样本3D CT VLMs的无监督多标签自适应。本文研究TTA对零样本3D CT VLMs的作用。控制诊断分析表明TTA是有条件的。随后引入CARVE方法,在多标签、三类和二元CT任务中,当基础模型已有判别能力时,CARVE能带来最一致的改进。这些结果确立了零样本3D CT VLMs的多标签TTA作为一个独特问题,以及CARVE作为一种基数感知解决方案。
英文摘要
3D CT vision-language models (VLMs) classify abnormalities from text prompts in a zero-shot manner, enabling cross-institution deployment where labels are scarce and clinical tasks shift faster than supervised models can be retrained. A real CT scan, however, typically contains several co-occurring abnormalities, and the reliability of zero-shot multi-label prediction under distribution shift remains poorly understood. Test-time adaptation (TTA) updates a model on unlabeled target scans without source data or target annotations, yet existing TTA methods target multi-class softmax prediction on natural images or 2D medical segmentation, and none addresses unsupervised multi-label adaptation for zero-shot 3D CT VLMs. We study when TTA helps zero-shot 3D CT VLMs. A controlled diagnostic analysis shows that TTA is conditional: the volumetric input must preserve the encoder's depth structure, and the base representation must transfer to the target cohort, with depth reduction alone lowering internal AUROC by more than 0.12. We then focus on the regime where the base model already separates present from absent abnormalities. We introduce CARVE (Cardinality-Aware Retained-View Entropy), the first TTA method for this setting. CARVE estimates a sample-specific positive-label cardinality $\hat{k}$, optimizes a top-$\hat{k}$ objective to preserve co-occurring abnormalities, and performs memory-efficient multi-view adaptation by scoring weak 3D views without gradients before updating on a retained subset. Across contrastive CT-CLIP and anatomy-aware fVLM, CARVE provides the most consistent improvements across multi-label, three-class, and binary CT tasks when the base model is already discriminative. These results establish multi-label TTA for zero-shot 3D CT VLMs as a distinct problem and CARVE as a cardinality-aware solution.
发表机构
- University of British Columbia(英属哥伦比亚大学)
- Vector Institute(向量研究所)
机构由 AI 辅助整理,请以论文原文为准。