arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

错误的器官,正确的物理:将超声心动图预训练迁移至肺超声用于结核病筛查

Wrong Organ, Right Physics: Transferring Echocardiography Pretraining to Lung Ultrasound for Tuberculosis Screening

Christiaan M. Geldenhuys, Joshua M. Jansen van Vüren, Véronique Suttels, Trevor Brokowski, Ablo P. Wachinou, Mary-Anne Hartley, Rensu P. Theart, Grant Theron, Thomas R. Niesler

arXiv 2610.03290首次发表:更新:

发表机构

University of Stellenbosch; Swiss Federal Institute of Technology (EPFL); National Teaching Hospital for Tuberculosis and Pulmonary Diseases (CNHU-PPC); South African Medical Research Council; Stellenbosch University(斯泰伦博斯大学; 瑞士联邦理工学院(EPFL); 国家结核病与肺病教学医院(CNHU-PPC); 南非医学研究理事会; 斯泰伦博斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究将超声心动图预训练迁移至肺超声结核筛查,发现编码器选择影响甚微,特征标准化显著提升性能,且标记数据规模是主要瓶颈。

AI 中文摘要

肺超声(LUS)在初级保健层面的结核病(TB)筛查中具有吸引力,但标记队列规模较小。超声心动图则没有此类限制,同时与肺超声共享相同的底层超声成像物理、信号处理及B模式外观。我们探究在预训练于高资源超声领域的编码器是否携带了在低资源领域中仍然可用的表征。仅编码器有所不同,涵盖三个架构家族的十七种编码器。其中,预训练于通用视频的潜在预测视频编码器(V-JEPA2-L)及其超声心动图对应版本(EchoJEPA-L)仅在预训练语料库上存在差异。这些编码器之间的选择并未决定分类结果,整个家族在2.71的测量分辨率下跨度达到2.50个百分点。真正影响任务的是特征条件化。在编码器与分类器之间对特征进行标准化处理,使全部十七种编码器平均提升了+1.23个百分点,p值为1.5×10^-5。在保留测试集上,基于开发折选出的每种编码器均比基线系统高出最多+2.57个百分点的受试者工作特征曲线下面积(AUROC),且在90%灵敏度下的特异性达到79.3%,而基线为60.3%。预先设定的对比,即EchoJEPA-L对V-JEPA2-L,测得差异为-0.16个百分点,p值为0.926。因此,我们没有发现共享超声物理本身使超声心动图成为比通用视频更具生产力的预训练语料库的证据,任何优势(若存在)均小于该队列所能分辨的范围。然而,视频编码器接收的是复制的静态图像,因此这种无效应现象究竟是反映预训练领域还是应用于静态帧的视频编码器,无法区分。限制因素在于标记队列而非编码器。

英文摘要

Lung ultrasound (LUS) is attractive for tuberculosis (TB) screening at primary-care level, but labelled cohorts are small. Echocardiography carries no such constraint, while sharing the same underlying ultrasound imaging physics, signal processing and B-mode appearance as LUS. We ask whether an encoder pretrained on that high-resource ultrasound domain carries representations that remain usable in the low-resource one. Only the encoder varies, across seventeen encoders spanning three architecture families. Among them, a latent-predictive video encoder pretrained on generic video (V-JEPA2-L) and its echocardiography counterpart (EchoJEPA-L) differ in pretraining corpus alone. The choice among these encoders does not resolve the classification, the whole family spanning 2.50 percentage points against a measurement resolution of 2.71. What moves the task instead is feature conditioning. Standardising the features between the encoder and the classifier improves all seventeen encoders by a mean of +1.23 percentage points at $p=1.5\times10^{-5}$. On the held-out test set every encoder selected on the development folds stands above the baseline system by up to +2.57 percentage points of area under the receiver operating characteristic curve (AUROC), and specificity at 90% sensitivity reaches 79.3% against 60.3%. The contrast specified in advance, EchoJEPA-L against V-JEPA2-L, measures -0.16 percentage points at $p=0.926$. We therefore find no evidence that shared ultrasonic physics alone makes echocardiography a more productive pretraining corpus than generic video, and any advantage, if present, is smaller than this cohort can resolve. The video encoders receive replicated still images, however, so whether this absence of an effect reflects the pretraining domain or a video encoder applied to static frames cannot be separated. The limiting factor is the labelled cohort rather than the encoder.

Comments10 pages, 3 figures, 4 tables. Accepted at SATNAC 2026, Drakensberg, South Africa, 11-14 October 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑