arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16551cs.CV

哪种预训练任务可迁移?肺超声的自监督预训练目标

Which Pretext Task Transfers? Self-Supervised Pretraining Objectives for Lung Ultrasound

Moein Heidari, Junbo Rao, Jai Choraria, Wenjin Chen, David J. Foran, Ilker Hacihaliloglu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在肺超声上比较对比学习、掩码重建和JEPA三类自监督目标,发现POCUS上VideoMAE和V-JEPA优于MoCo,但跨数据集时排名反转,MoCo迁移最佳,表明单一数据集探针准确率不足以判断迁移性能。

中文摘要 AI 辅助

自监督学习(SSL)可以减少对标注医学图像的需求,但对于肺超声(LUS)而言,预训练目标的选择仍不明确。对比学习、掩码重建和联合嵌入预测架构(JEPA)在目标定义的空间上有所不同,但现有的超声研究在语料库、骨干网络和评估协议不同的情况下进行比较。我们使用相同的编码器骨干网络、预训练语料库、优化调度和冻结评估协议,比较了这三类目标。编码器在COVID-BLUeS LUS视频上预训练,并在5%、10%、50%和100%标签预算下,使用线性、k近邻($k$NN)和注意力探针进行评估。评估在POCUS上进行,采用患者级五折交叉验证,并在独立获取的Mendeley-Uganda数据集上进行,该数据集被排除在预训练和探针拟合之外。在全部标签预算下,使用线性探针时,VideoMAE和V-JEPA在POCUS上分别达到$66.5 \pm 13.1$和$65.4 \pm 11.7$的平衡准确率,而MoCo达到$42.1 \pm 1.2$。在Mendeley-Uganda上,排名反转:MoCo表现最佳,为$62.7 \pm 1.0$,其次是VideoMAE的$53.8 \pm 2.8$,而V-JEPA接近随机水平,为$35.1 \pm 4.9$。这些结果表明,仅凭POCUS探针准确率无法确定跨数据集迁移最佳的目标。我们还概述了计划中的表示级分析,以检查这种反转。代码可在以下https URL公开获取。

英文摘要

Self-supervised learning (SSL) can reduce the need for labelled medical images, but the choice of pretext objective remains unclear for lung ultrasound (LUS). Contrastive learning, masked reconstruction, and joint-embedding predictive architectures (JEPA) differ in the space in which their targets are defined, yet existing ultrasound studies compare them under different corpora, backbones, and evaluation protocols. We compare these three objective families using the same encoder backbone, pretraining corpus, optimisation schedule, and frozen-evaluation protocol. Encoders are pretrained on COVID-BLUeS LUS videos and evaluated with linear, $k$NN, and attentive probes at 5\%, 10\%, 50\%, and 100\% label budgets. Evaluation is performed on POCUS using patient-level five-fold cross-validation and on the independently acquired Mendeley-Uganda dataset, which is excluded from both pretraining and probe fitting. At the full label budget under linear probing, VideoMAE and V-JEPA achieve $66.5 \pm 13.1$ and $65.4 \pm 11.7$ balanced accuracy on POCUS, while MoCo achieves $42.1 \pm 1.2$. On Mendeley-Uganda, the ranking reverses: MoCo performs best at $62.7 \pm 1.0$, followed by VideoMAE at $53.8 \pm 2.8$, while V-JEPA falls near chance at $35.1 \pm 4.9$. These results show that POCUS probe accuracy alone does not identify the objective that transfers best across datasets. We also outline planned representation-level analyses to examine this reversal. Code is publicly available at https://github.com/moeinheidari7829/LUSVideoSSL.

发表机构

  • University of British Columbia(不列颠哥伦比亚大学)
  • Rutgers Cancer Institute of New Jersey(新泽西州罗格斯癌症研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑