发表机构
The Chinese University of Hong Kong; The Hong Kong Polytechnic University(香港中文大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对手术三元组识别的三大冲突问题,提出MoeCo框架,通过组件定制适配器、协同梯度学习及知识驱动混合专家机制,在CholecT45和CholecT50数据集上验证了方法的有效性与优越性。
AI 中文摘要
手术动作三元组识别是上下文感知机器人辅助手术中的关键任务,通过识别器械、动作动词、目标及其关联,实现自动手术动作感知。然而,现有研究因三大问题难以分析此类复杂手术场景:(1)特征空间纠缠引发的组件级优化冲突;(2)数据严重不平衡导致的类别级优化冲突;(3)缺乏领域知识指导,限制了模型的可解释性与鲁棒性。为应对这些挑战,本文提出一种由知识驱动学习支撑的混合专家引导协同优化(MoeCo)框架。在协同优化流程中,为缓解组件级冲突,引入组件定制适配器,解耦时空域内的任务特定特征,促进组件有效专业化;接着,开发协同梯度学习策略处理类别级冲突,自适应平衡正负梯度以增强稀有类别的感知。值得注意的是,受外科领域专业知识启发,引入知识驱动的混合专家机制,通过激活的专家动态整合多模态大语言模型引导的知识,为协同优化流程补充更具表达力和鲁棒性的表征。在公开的CholecT45和CholecT50数据集上开展的大量实验,证实了所提协同优化流程的有效性,以及通过知识驱动的混合专家机制整合动态先验的优越性。
英文摘要
Surgical action triplet recognition constitutes a critical task in context-aware robot-assisted surgery, facilitating automatic surgical action perception by identifying instrument, verb, target, and their association. However, existing works struggle to analyze such complex surgical scenes due to three main issues: (1) component-level optimization conflicts caused by entangled feature spaces, (2) category-level optimization conflicts arising from severe data imbalance, and (3) lack of domain knowledge guidance that limits model interpretability and robustness. To address these challenges, we propose a Mixture-of-Experts-guided Co-Optimization (\textit{MoeCo}) framework powered by knowledge-driven learning. Within the co-optimization pipeline, to first mitigate component-level conflicts, we introduce a component-tailored adapter that disentangles task-specific features across spatial-temporal regimes, facilitating effective component specialization. Next, we develop a coordinated gradient learning strategy to handle category-level conflicts, which adaptively rebalances positive-negative gradients to enhance the perception of rare categories. Notably, inspired by surgical domain expertise, we introduce a knowledge-driven mixture-of-experts mechanism that dynamically integrates multimodal large language model-guided knowledge via activated experts, thereby enriching the co-optimization pipeline with more expressive and robust representations. Extensive experiments on the public CholecT45 and CholecT50 datasets confirm the effectiveness of the proposed co-optimization pipeline and the superiority of dynamic priors integration via the knowledge-driven mixture-of-experts mechanism.
CommentsAccepted in TMI