arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于梯度的潜在分解揭示了弱监督乳腺钼靶成像中特征退化的机制

Gradient-Based Latent Decomposition Reveals Mechanisms of Feature Degradation in Weakly Supervised Mammography

Vinceline Bertrand, Ionut Cardei

arXiv 2607.24835首次发表:更新:

发表机构

Florida Atlantic University(佛罗里达大西洋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究弱监督乳腺钼靶成像中特征退化问题,引入基于梯度的正交潜在分解用于分层变分自编码器,划分潜在空间,通过实验得出模型相关指标,经潜在消融等验证,揭示高维空间中粗监督信号对细粒度特征的影响。

AI 中文摘要

弱监督分层模型存在持续的不对称性:粗病变类型特征在重建时得以保留,而细粒度恶性线索会退化,这对乳腺癌筛查管道的临床可靠性有直接影响。我们引入基于梯度的正交潜在分解用于分层变分自编码器(H-VAEs),以机械地解释这种不对称性。潜在空间被划分为由粗监督梯度塑造的任务对齐组件($z_1$)和捕获剩余表示能力的正交残差($z_{\text{res}}$)。在来自CBIS-DDSM的3550个乳腺钼靶感兴趣区域(ROIs)上,只有约4.4%的潜在幅度与监督梯度对齐,约95.6%在正交残差中,细粒度病理预测主要依赖于此。该模型实现了第一阶段AUC为0.866,第二阶段AUC为0.552,重建稳定性差距为$\Delta_{\text{diag}} = 5\%$($p = 0.005$),分类差距为$\Delta_{\text{AUC}} = 0.314$($p < 0.001$)。潜在消融证实两个任务的特征都大量存在于$z_{\text{res}}$中,从结构上解释了为什么重建会不成比例地降低病理稳定性。与多实例学习(MIL)和多任务学习(MTL)的比较证实了跨架构和模态的泛化。这些发现表明,在高维空间中,单个粗监督信号仅隔离出一个稀疏的一维潜在方向,迫使关键的细粒度特征进入易受影响的残差子空间。

英文摘要

Weakly supervised hierarchical models exhibit a persistent asymmetry: coarse lesion-type features are preserved under reconstruction while fine-grained malignancy cues degrade---a pattern with direct consequences for the clinical reliability of breast cancer screening pipelines. We introduce gradient-based orthogonal latent decomposition for hierarchical Variational Autoencoders~(H-VAEs) to mechanistically explain this asymmetry. The latent space is partitioned into a task-aligned component~($z_1$), shaped by coarse supervisory gradients, and an orthogonal residual~($z_{\text{res}}$) capturing remaining representational capacity. On~3,550 mammographic Regions of Interest~(ROIs) from CBIS-DDSM, only~$\sim$4.4\% of latent magnitude aligns with supervisory gradients, leaving~$\sim$95.6\% in the orthogonal residual upon which fine-grained pathology prediction primarily depends. The model achieves Stage-1~AUC~0.866 and Stage 2~AUC~0.552, with a reconstruction stability gap of $Δ_{\text{diag}}=5\%$ ($p=0.005$) and a classification gap of $Δ_{\text{AUC}}=0.314$ ($p{<}0.001$). Latent ablation confirms that features for both tasks reside heavily in~$z_{\text{res}}$, structurally explaining why reconstruction degrades pathology stability disproportionately. Comparisons with Multi-Instance Learning~(MIL) and Multi-Task Learning~(MTL) confirm generalization across architectures and modalities. These findings reveal that in high-dimensional spaces, a single coarse supervisory signal isolates only a sparse 1D latent direction, forcing critical fine-grained features into the vulnerable residual subspace.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑