arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于基础模型的高效数据采样(FEEDS):用于泛癌、多示踪剂PET/CT数据集的标签高效训练策略

Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets

Biratal Raj Wagle, Bashirul Azam Biswas, Grant Chau, Matthew E. Maeder, Muhammad Azeem Arshad, Michael S. Leapman, James B. Yu, Indrani Bhattacharya

arXiv 2608.11076首次发表:更新:

发表机构

Geisel School of Medicine at Dartmouth; Dartmouth Hitchcock Medical Center; Yale University(达特茅斯盖泽尔医学院; 达特茅斯-希区柯克医疗中心; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出FEEDS策略,利用视觉基础模型嵌入选择信息性和多样性未标注病例,在减少70%标注负担时达到全标注训练性能,解决了PET/CT病灶分割的标签稀缺问题。

AI 中文摘要

全身PET/CT成像中的自动化病灶分割可辅助临床医生在不同放射性示踪剂和癌症类型下开展癌症检测、分期及治疗规划。然而,训练能捕捉病灶大小、分布及外观差异的病灶分割模型需要大量标注数据集,其创建既耗时又依赖专业知识。因此,在有限标注PET/CT数据上训练的模型往往缺乏临床应用所需的准确性和泛化能力。我们提出FEEDS(Foundation model-Enabled Efficient Data Sampling,基于基础模型的高效数据采样),这是一种标签和计算高效的学习策略,利用视觉基础模型的嵌入来选择最具信息性和多样性的未标注病例供专家标注。与无监督、半监督及主动学习方法不同,FEEDS是一种仅需有限、代表性训练集的单步训练范式,兼具标签和计算高效性。我们使用AutoPET-III数据集对FEEDS进行训练和验证,在三个保留测试集(AutoPET-III、DeepPSMA及达特茅斯-希区柯克医学中心的内部数据集)上测试其准确性和泛化性,在体素、病灶及解剖区域层面评估临床效用,以评估高危区域的性能和治疗规划效用。FEEDS的表现优于基于随机采样的标注、基于伪标签的半监督学习以及仅用有限标注数据的训练,可在所有三个测试集、FDG和PSMA示踪剂及多种疾病间实现泛化,在减少70%标注负担的情况下达到了100%全标注训练的性能。FEEDS通过提供一种从大型未标注临床库中构建具有代表性和多样性的标注队列的实用方法,解决了自动病灶分割框架中的标签稀缺问题。

英文摘要

Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotracers and cancer types. However, training lesion segmentation models that capture variations in lesion size, distribution, and appearance requires large annotated datasets, whose creation is both time- and expertise-intensive. As a result, models trained on limited labeled PET/CT data often lack the accuracy and generalizability needed for clinical use. We present FEEDS (Foundation model-Enabled Efficient Data Sampling), a label- and compute-efficient learning strategy that uses vision foundation model embeddings to select the most informative and diverse unlabeled cases for expert annotation. Unlike unsupervised, semi-supervised, and active learning approaches, FEEDS is a one-step training paradigm requiring only a limited, representative training set, making it label- and compute-efficient. We train and validate FEEDS using the AutoPET-III dataset. We test its accuracy and generalizability on three held-out sets: AutoPET-III, DeepPSMA, and an internal Dartmouth-Hitchcock Medical Center dataset. We evaluate clinical utility at the voxel, lesion, and anatomic region level to assess performance in high-risk areas and treatment planning utility. FEEDS outperforms random-sampling-based labeling, pseudolabel-based semi-supervised learning, and training with limited labeled data alone. It generalizes across all three test sets, FDG and PSMA tracers, and multiple diseases, matching fully-labeled (100\%) training performance with 70\% less annotation burden. FEEDS addresses the challenge of label scarcity in an automatic lesion segmentation framework by providing a practical approach for constructing representative and diverse annotation queues from large, unannotated clinical repositories.

CommentsCode is publicly available on https://github.com/Image-and-Multimodal-Data-Analytics/FEEDS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑