AI 中文总结
研究利用超声 IQ 数据估计声速和衰减的难题,提出 IQ-JEPA 架构,先无标签预训练编码器预测 IQ 区域潜在表示,再微调。实验表明该方法提高标签效率,自监督是主导因素,为定量超声基础模型迈出第一步。
AI 中文摘要
组织中的声速是聚焦成像的前提且具有诊断价值,但从原始脉冲回波通道数据恢复声速本质上是一个非线性逆问题。有监督求解器速度快但依赖标签,模拟声速标签成本高,而大量真实通道数据无标签。我们提出 IQ-JEPA 来利用这两种数据类型。先对编码器进行无标签预训练,从可见上下文预测掩码同相和正交(IQ)区域的潜在表示,再在模拟图上微调。声速在 IQ 信号中表现为相位差,对恒定相位偏移不变。编码器是直接对复信号操作的厄米特视觉变换器,其注意力对该相位等变,共轭积前馈对其不变,所以编码器读取类似于经典相干方法使用的量。在 2.5MHz 的 79293 次全波 2.5 模拟中,在 63435 次无标签采集上预训练,在 10000 个标签时达到 15.60m/s 的精度。这比监督训练的标签效率提高了约三倍,在 1000 个标签时超过四倍。比 InversionNet 基线低约 2.2 倍,全标签时为 8.71m/s。随着更多无标签预训练数据,增益仍在增加。我们的比较表明自监督是主导因素。相同的编码器可迁移,其冻结特征能揭示声速和衰减,分层和腹部体模之间进行交叉分布预训练对精度影响不大。我们将此视为迈向定量超声基础模型的第一步。
英文摘要
The speed of sound in tissue is a prerequisite for well-focused imaging and has diagnostic value, but recovering it from raw pulse-echo channel data is fundamentally a nonlinear inverse problem. Learned solvers are fast yet label hungry. Simulated sound-speed labels are expensive, while abundant real channel data is unlabeled. We propose IQ-JEPA to exploit both data types. An encoder is pretrained without labels to predict the latent representation of masked in-phase and quadrature (IQ) regions from visible context, then fine-tuned on simulated maps. Sound speed appears in the IQ signal as a phase difference, invariant to the constant phase offset. The encoder is a Hermitian vision transformer that operates on the complex signal directly. Its attention is equivariant to that phase and its conjugate-product feed-forward is invariant to it, so the encoder reads a quantity analogous to the one classical coherence methods use. On 79,293 Fullwave 2.5 simulations at 2.5 MHz, pretraining on the 63,435 unlabeled acquisitions reaches 15.60 m/s at 10,000 labels. This is a roughly threefold gain in label efficiency over supervised training, growing to over fourfold at 1,000 labels. It is about 2.2x below an InversionNet baseline, and 8.71 m/s at full labels. The gain still grows with more unlabeled pretraining data. Our comparisons point to self-supervision as the dominant factor. The same encoder transfers. Its frozen features expose sound speed and attenuation, and cross-distribution pretraining between layered and abdominal phantoms costs little accuracy. We see this as a first step toward a foundation model for quantitative ultrasound.