LaSeD:面向纯视觉手术阶段识别的标签语义自蒸馏
LaSeD: Label-Semantic Self-Distillation for Visual-Only Surgical Phase Recognition
浏览论文内容
中文总结 AI 辅助
LaSeD利用阶段名称作为特权上下文进行自蒸馏,在纯视觉推理下提升手术阶段识别性能,显著超越基线。
中文摘要 AI 辅助
手术阶段识别将每个视频帧映射到具有临床意义的工作流阶段,支持上下文感知的辅助、文档记录和术后分析。大多数方法仅将阶段标注视为类别ID,而近期的手术视觉-语言模型通常需要额外的视频-文本数据、字幕或指令微调。我们提出LaSeD,一种标签语义自蒸馏框架,利用阶段名称作为训练时的特权上下文,同时保持纯视觉部署,无需真实阶段名称提示。LaSeD从相同的预训练VLM检查点初始化一个冻结的教师模型和一个学生模型。教师模型接收帧、固定任务提示和真实阶段名称提示;学生模型接收相同的帧和提示但不含提示,且仅优化其视觉编码器。训练结合了硬阶段令牌监督和来自缓存教师表征的特征级蒸馏。在推理时,移除教师模型和提示,学生模型通过约束的数字令牌逻辑预测Cholec80的七个阶段之一,无需额外的分类头。在Cholec80评估分割上,LaSeD实现了86.20%的准确率、77.75%的宏召回率、78.13%的宏精确率和64.15%的宏Jaccard指数。在相同协议下,它分别将纯视觉的Qwen3-VL-4B基线提升了9.45、9.42、11.93和12.22个百分点。纯视觉意味着图像是唯一的样本特定推理输入,而所有帧共享相同的固定任务提示。这些结果表明,阶段名称为将VLM适应手术工作流分析提供了有用的低成本信号。(代码即将发布。)
英文摘要
Surgical phase recognition maps each video frame to a clinically meaningful workflow phase, supporting context-aware assistance, documentation, and postoperative analysis. Most methods treat phase annotations only as class IDs, whereas recent surgical vision-language models often require additional video--text data, captions, or instruction tuning. We propose \emph{LaSeD}, a label-semantic self-distillation framework that uses phase names as privileged training-time context while retaining visual-only deployment without a ground-truth phase-name hint. LaSeD initializes a frozen teacher and a student from the same pretrained VLM checkpoint. The teacher receives the frame, a fixed task prompt, and the ground-truth phase-name hint; the student receives the same frame and prompt without the hint, and only its visual encoder is optimized. Training combines hard phase-token supervision with feature-level distillation from cached teacher representations. At inference, the teacher and hint are removed, and the student predicts one of the seven Cholec80 phases through constrained digit-token logits without an additional classifier head. On the Cholec80 evaluation split, LaSeD achieves 86.20\% accuracy, 77.75\% macro recall, 78.13\% macro precision, and 64.15\% macro Jaccard. Under the identical protocol, it improves a visual-only Qwen3-VL-4B baseline by 9.45, 9.42, 11.93, and 12.22 percentage points, respectively. Visual-only means that the image is the only sample-specific inference input, while all frames share the same fixed task prompt. These results suggest that phase names provide a useful low-cost signal for adapting VLMs to surgical workflow analysis. (The code will be published soon.)
发表机构
- Heidelberg University(海德堡大学)
- University Medical Center Mannheim(曼海姆大学医学中心)
- Bielefeld University(比勒费尔德大学)
机构由 AI 辅助整理,请以论文原文为准。