DINOv3-MIL:基于KiTS23数据集基础模型补丁令牌的肾脏多标签肿瘤和囊肿检测
DINOv3-MIL: Per-Kidney Multi-Label Tumour and Cyst Detection from Foundation-Model Patch Tokens on KiTS23
浏览论文内容
中文总结 AI 辅助
研究在KiTS23数据集上,基于冻结的DINOv3 ViT-H/16+特征,比较CLS令牌线性探针、门控注意力MIL和原型头三种聚合器用于肾肿瘤/囊肿检测,发现注意力MIL效果最佳,还揭示了原型头在囊肿检测中可解释性与性能的权衡。
中文摘要 AI 辅助
在自然图像上训练的基础视觉模型无需进行领域预训练即可转移到医学任务中,但体积分类需要为每个研究聚合数万个补丁令牌,并且聚合器限制了所得模型的可解释性。我们在KiTS23数据集(966个肾脏;n=97个测试样本)上,针对相同的冻结DINOv3 ViT-H/16+特征比较了三种聚合器用于肾肿瘤/囊肿检测:CLS令牌线性探针、对55296个补丁令牌进行门控注意力多实例学习(MIL)以及遵循ProtoViT的原型头。注意力MIL在肿瘤检测(AUROC为0.74,95%置信区间0.64 - 0.83)和囊肿检测(AUROC为0.80,0.70 - 0.88)方面取得了最高的AUROC,在标注病变内注意力富集比随机情况高7.5 - 9.8倍。原型头在囊肿检测中表现不佳(AUROC为0.51),揭示了在此令牌规模下可解释性与性能之间的权衡。
英文摘要
Foundation vision models trained on natural images transfer to medical tasks without domain pre-training, but volumetric classification requires aggregating tens of thousands of patch tokens per study, and the aggregator constrains how the resulting model can be interpreted. We compare three aggregators on identical frozen DINOv3 ViT-H/16+ features for renal tumour/cyst detection on KiTS23 (966 kidneys; n=97 test): a CLS-token linear probe, gated attention multiple instance learning (MIL) over 55,296 patch tokens, and a prototype head following ProtoViT. Attention MIL achieves the highest AUROC for tumour (0.74, 95% CI 0.64-0.83) and cyst (0.80, 0.70-0.88), with attention enriched 7.5-9.8x over chance within annotated lesions. The prototype head does not transfer to cyst detection (AUROC 0.51), exposing an interpretability-performance trade-off at this token scale.