用于零样本手术阶段识别的大小模型协作
Large-Small Model Collaboration for Zero-Shot Surgical Phase Recognition
- The Chinese University of Hong Kong(香港中文大学)
- The Hong Kong Polytechnic University(香港理工大学)
- Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学医学智能与XR研究所)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究提出LaST大小模型协作框架,结合轻量模型时间建模能力与基础模型迁移性,实现零样本手术阶段识别,在未见领域自适应上性能显著优于基线及部分先进方法。
中文摘要 AI 辅助
手术阶段识别的特定任务轻量模型擅长捕捉时间动态,但在领域偏移下泛化能力较差;相反,手术基础模型(FMs)通过大规模预训练具备优异的迁移性,但缺乏显式时间建模往往会产生时间不一致的预测,导致性能下降。为利用两种范式的互补优势,我们提出了一种新颖的大小时间自适应框架LaST(Large-Small Temporal adaptation),该框架可实现对未见临床领域的零样本自适应。在LaST中,基础模型启动流程,生成帧级阶段先验作为初始弱监督;为有效利用这些含噪的阶段先验,我们引入了迭代时间优化方案,该方案整合了动态质量控制以过滤可靠预测,以及双模型交叉学习以缓解确认偏差;同时,轻量模型利用其内在的时间建模能力,逐步修正不一致的预测并在迭代过程中提升整体准确率;最后,采用循环重放策略形成闭环:经过优化、更准确的预测被用作后续迭代的升级监督信号,促进标签质量与模型能力的自增强演化。大量实验表明,LaST在零样本手术阶段识别中对未见领域实现了鲁棒自适应,在准确率上较基线方法PeskaVLP提升了24.85%-43.17%,甚至超过了全监督线性探测及若干最先进的少样本方法。代码将发布在该httpsURL。
英文摘要
Task-specific lightweight models for surgical phase recognition excel at capturing temporal dynamics but generalize poorly under domain shift. Conversely, surgical foundation models (FMs) offer superior transferability via large-scale pretraining, yet their lack of explicit temporal modeling often yields temporally inconsistent predictions, leading to degraded performance. To exploit the complementary strengths of both paradigms, we propose \textbf{La}rge-\textbf{S}mall \textbf{T}emporal adaptation (\textbf{LaST}), a novel large-small collaborative framework that enables zero-shot adaptation to unseen clinical domains. In LaST, the FM initiates the pipeline by generating frame-level phase priors that serve as initial weak supervision. To effectively utilize these noisy phase priors, we introduce an iterative temporal refinement scheme that integrates dynamic quality control to filter reliable predictions and dual-model cross-learning to mitigate confirmation bias. Simultaneously, the lightweight model leverages its intrinsic temporal modeling ability to progressively correct inconsistent predictions and enhance overall accuracy across iterations. At the end, a cycle replay strategy is employed to close the loop: the refined, more accurate predictions are utilized as upgraded supervision signals for the subsequent iterations, fostering a self-reinforcing evolution of both label quality and model capability. Extensive experiments demonstrate that LaST achieves robust adaptation to unseen domains for zero-shot surgical phase recognition, outperforming the baseline (PeskaVLP) by 24.85\%-43.17\% in accuracy and even surpassing fully supervised linear probing and several state-of-the-art few-shot approaches. Codes will be released at https://github.com/YIYIZH/LaST.