发表机构
City University of Hong Kong (Dongguan); The Hong Kong Polytechnic University(香港城市大学(东莞); 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对医学影像分析深度学习模型的灾难性遗忘问题,提出STAIL框架,通过SCB和LSAM方法,在三个异构医学数据集上提升基线性能,即插即用且效果显著。
AI 中文摘要
应用于医学影像分析的深度学习模型在动态环境中持续适配新临床任务时,会遭受严重的灾难性遗忘问题。主流增量学习方法通常通过回放原始历史图像来缓解该问题,但这种像素级回放会产生显著的存储开销,引发隐私担忧,且稀疏样本无法充分捕捉真实数据分布。受人类认知机制启发,我们提出了一种名为语义文本锚定增量学习(Semantic Text-Anchored Incremental Learning, STAIL)的新型框架,用于处理序列临床任务。为克服回放瓶颈,STAIL引入了非对称语义整合缓冲器(Semantic Consolidation Buffer, SCB),通过结合少量图像锚点与大量文本描述,以极低存储成本实现旧任务的密集语义重建。此外,我们设计了一种源自大语言模型(Large Language Model, LLM)的语义锚定机制(Semantic Anchoring Mechanism, LSAM),利用冻结大语言模型的稳定语义空间作为发展先验,该机制将演化的视觉特征明确锚定到文本表示,在宏观和微观层面引导并约束可塑性与稳定性。在涵盖眼底、超声和X射线成像的三个异构医学数据集上开展的大量实验表明,STAIL是一种极具效的即插即用模块,它全面提升了多种现有基线方法的性能,在持续性能指标AAA-AUC上平均提升2.24%,在减少遗忘的指标BWT-AUC上平均提升3.55%。代码已开源。
英文摘要
Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dynamic environments. Mainstream incremental learning methods typically mitigate this by rehearsing raw historical images. However, this pixel-level rehearsal incurs significant storage overhead, raises privacy concerns, and fails to adequately capture the true data distribution with sparse exemplars. Inspired by human cognitive mechanisms, we propose a novel framework termed Semantic Text-Anchored Incremental Learning (STAIL) for sequential clinical tasks. To overcome the rehearsal bottleneck, STAIL introduces an asymmetric semantic consolidation buffer (SCB). By incorporating a minimal set of image anchors and extensive textual descriptions, the SCB enables dense semantic reconstruction of old tasks at a minimal storage cost. Furthermore, we design an LLM-derived Semantic Anchoring Mechanism (LSAM) that leverages the stable semantic space of frozen large language models as developmental priors. This mechanism explicitly anchors evolving visual features to textual representations, guiding and constraining plasticity and stability at both macroscopic and microscopic levels. Extensive experiments across three heterogeneous medical datasets, covering fundus, ultrasound, and X-ray imaging, demonstrate that STAIL acts as a highly effective plug-and-play module. It comprehensively enhances the performance of various existing baselines, achieving average gains of 2.24\% in AAA-AUC for sustained performance and 3.55\% in BWT-AUC for reduced forgetting. Code is available.