arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15268cs.CV

用于超声心动图心肌梗死定位的运动条件多视图融合

Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography

  • University of Oxford(牛津大学)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Guang Yang, Wentian Xu, Siyu Wang, Betty Raman, Lei Li, Vicente Grau

AI总结:

针对超声心动图心肌梗死定位问题,提出MCF-Net框架,融合心肌运动线索与基础模型表示,通过极稀疏监督建模心脏运动,利用运动衍生软掩码提供先验,跨视图整合运动和视觉优化预测,在节段级定位上性能优于现有方法。

AI中文摘要:

心肌梗死是全球主要死因。超声心动图是评估心肌梗死的常用方法,局部室壁运动异常是关键指标。以往基于学习的心肌运动分析方法存在局限性。基础模型改进了基于视觉的超声心动图分析,但多数方法基于单视图,在视图依赖模糊性下,尤其是心尖视图,节段级定位不可靠。为此提出MCF-Net,一种运动引导的多视图融合框架,融合心肌运动线索与基础模型表示来定位梗死。使用EchoPrime提取视觉特征,通过极稀疏监督建模心脏运动,运动衍生的段感知软掩码提供空间先验,运动条件融合机制跨视图整合运动和视觉来优化预测。在节段级心肌梗死定位上,MCF-Net取得了72.4%的F1值和84.9%的准确率,优于现有方法。

英文摘要:

Myocardial infarction (MI) remains a leading cause of mortality worldwide. Echocardiography (Echo) is a widely available modality for MI assessment, where regional wall motion abnormality is a key indicator. Prior learning based methods for myocardial motion analysis often use handcrafted descriptors or densely supervised estimation, but the need for extensive annotation limits applicability. Foundation models have recently improved vision-based Echo analysis; however, most methods operate on single views and segment-level localization remains unreliable under view-dependent ambiguity, especially in apical views. To address this, we propose MCF-Net, a novel motion-guided multi-view fusion framework that fuses myocardial motion cues with foundation model representations to localize infarction. Visual features are extracted using EchoPrime, a pretrained Echo foundation model shared across dual views. Cardiac motion is modeled with extremely sparse supervision: a single annotated template frame is transferred across videos to initialize point tracking, avoiding dense labels. Motion-derived segment-aware soft masks provide coarse spatial priors that selectively enhance features for challenging myocardial segments. A motion-conditioned fusion mechanism then integrates motion and vision across views, refining predictions without overriding strong appearance cues. On segment-level MI localization, MCF-Net achieves 72.4\% F1 and 84.9\% accuracy, outperforming state-of-the-art motion-only, vision-only, and fusion baselines.

↑