基于解剖学引导的基础模型适配与病例内原型监督的胎儿超声盲扫标准平面检测
Anatomy-Guided Foundation Model Adaptation with Within-Case Prototype Supervision for Standard Plane Detection in Fetal Ultrasound Blind Sweeps
浏览论文内容
中文总结 AI 辅助
针对胎儿超声盲扫中阳性帧占比低的不平衡问题,提出AnatoProto框架,通过四个组件适配BiomedCLIP,在ACOUSLIC-AI基准上F1值优于现有基线
中文摘要 AI 辅助
在低成本产科盲扫中检测胎儿腹围标准平面是一个高度不平衡的帧分类问题:阳性帧在序列中占比不足3%,形成短的连续片段,且难以被现成的超声和视觉基础模型处理。我们提出AnatoProto,一种轻量级序列级框架,通过四个组件适配冻结的BiomedCLIP编码器以用于胎儿盲扫:(i)解剖学加权空间池化,使用nnU-Net的腹部区域概率作为空间先验,对BiomedCLIP的补丁标记进行重新加权,从而将冻结的语义特征聚合到具有解剖学意义的区域;(ii)病例内原型损失,将每个帧嵌入拉向同一次扫查中阳性帧的均值,利用帧级别不可用的病例级结构;(iii)三级级联细化(帧→片段→病例级弃权(不执行)器),将预测单元从噪声帧提升到受结构约束的片段;(iv)混合预测头,联合建模每帧稳定性和帧间边界转换以抑制边界误报。在ACOUSLIC-AI基准上,AnatoProto达到测试F1值67.72,比最强的基础模型基线(FetalCLIP + PRS,F1=54.52)高出13.20个F1,比最强的视频时间动作检测基线(TriDet + PRS)高出15.76个F1。由嵌入几何和配对自助法置信区间支持的协同研究表明,原型损失和解剖学加权空间池化并非可加关系:单独应用原型损失会使召回率降低12个百分点,但与解剖学加权空间池化结合时会使召回率提高6.5个百分点——这种符号反转可追溯到病例内原型的准确性。
英文摘要
Detecting the fetal abdominal circumference standard plane in low-cost obstetric blind sweeps is a highly imbalanced frame-classification problem: positive frames account for under 3% of a sequence, form short contiguous segments, and are poorly handled by off-the-shelf ultrasound and vision foundation models. We propose AnatoProto, a lightweight sequence-level framework that adapts a frozen BiomedCLIP encoder to fetal blind sweeps through four components: (i) anatomy-weighted spatial pooling that uses nnU-Net abdominal-region probabilities as a spatial prior to reweight BiomedCLIP patch tokens, so frozen semantic features are aggregated onto anatomically meaningful regions; (ii) a within-case prototype loss that pulls each frame embedding toward the mean of positive frames of the same sweep, exploiting case-level structure unavailable at the frame level; (iii) a three-stage cascade refinement (frame->segment->case-level rejecter) that lifts the prediction unit from noisy frames to structurally-constrained segments; and (iv) a hybrid prediction head that jointly models per-frame stability and inter-frame boundary transitions to suppress boundary false positives. On the ACOUSLIC-AI benchmark, AnatoProto reaches a test F1 of 67.72, outperforming the strongest foundation-model baseline (FetalCLIP + PRS, F1 = 54.52) by +13.20 F1 and the strongest video temporal-action-detection baseline (TriDet + PRS) by +15.76 F1. A synergy study, backed by embedding geometry and paired-bootstrap confidence intervals, shows that the prototype loss and anatomy-weighted pooling are not additive: applied alone the prototype loss reduces recall by 12 points, but combined with anatomy-weighted pooling it increases recall by 6.5 points -- a sign-flip we trace to the accuracy of the within-case prototype.
发表机构
- School of Computing, University of Leeds(利兹大学计算学院)
机构由 AI 辅助整理,请以论文原文为准。