自监督预训练在数据稀缺时对视网膜疾病进展建模帮助最大
Self-supervised Pre-training Helps Retinal Disease Progression Modelling Most When Data Is Scarce
- Hertie Institute for AI in Brain Health, Faculty of Medicine, University of Tübingen(赫蒂人工智能脑健康研究所,图宾根大学医学院)
- Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对纵向数据稀缺问题,研究对比多种自监督预训练策略,发现冻结编码器加轻量生存头在少量标记样本下即可达到临床合理判别,且迁移由自监督目标主导。
AI中文摘要:
对疾病随时间进展进行建模需要纵向影像队列,这类队列稀缺且规模小,而横断面数据(每位参与者一张图像)则十分丰富。在此类数据上进行自监督预训练提供了一种弥合这一差距的方法,但尚不清楚哪种策略最能支持进展建模,以及该答案如何依赖于标记纵向数据的数量。我们针对年龄相关性黄斑变性(AMD)研究了这一问题,在大型横断面NAKO队列上预训练编码器,并在纵向AREDS数据集上预测晚期AMD的时间。我们比较了内部自监督编码器与通用基础模型(DINOv2)和领域专用基础模型(RETFound),涵盖对比学习、掩码自编码和自蒸馏目标,在冻结和微调协议下,以及从100到32,250个样本的标记训练集上进行了比较。哪种模型表现最佳取决于编码器的使用方式。当编码器被冻结且标签很少(纵向队列的典型情形)时,预训练表示仅需几百个标记样本即可达到临床合理的判别能力,而从头训练的模型则无法做到;这种优势在微调下会消失。迁移由自监督目标而非语料库规模或领域匹配决定,因此在适度的横断面队列上预训练的编码器能达到或超过规模大得多的领域内基础模型。总之,这些结果为在纵向数据稀缺时构建进展模型提供了实用方案:一个冻结的自监督编码器配合一个轻量级生存头。
英文摘要:
Modelling how a disease progresses over time requires longitudinal imaging cohorts, which are scarce and small, whereas cross-sectional data -- one image per participant -- is abundant. Self-supervised pre-training on such data offers a way to bridge this gap, but it is unclear which strategy best supports progression modelling, or how that answer depends on the amount of labelled longitudinal data. We study this for age-related macular degeneration (AMD), pre-training encoders on the large cross-sectional NAKO cohort and predicting time to late AMD on the longitudinal AREDS dataset. We compare in-house self-supervised encoders against a general-purpose (DINOv2) and a domain-specific (RETFound) foundation model, across contrastive, masked-autoencoding, and self-distillation objectives, under frozen and fine-tuned protocols, and across labelled training sets from 100 to 32,250 examples. Which model performs best depends on how the encoder is used. When the encoder is frozen and labels are few -- the regime typical of longitudinal cohorts -- pre-trained representations reach clinically reasonable discrimination from a few hundred labelled samples, while models trained from scratch do not; this advantage fades under fine-tuning. Transfer is governed by the self-supervision objective rather than corpus scale or domain match, so that an encoder pre-trained on a modest cross-sectional cohort matches or exceeds a far larger in-domain foundation model. Together, these results offer a practical recipe for building progression models where longitudinal data is scarce: a frozen self-supervised encoder with a lightweight survival head.