发表机构
The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出FLAIR框架,利用光流学习动作先验并注入潜动作表示,实现无参考条件的手术视频生成,同时构建数据集SurgActionClip-30K和评估指标SurgMetrics,提升生成视频的动作一致性与临床相关性。
AI 中文摘要
手术视频生成在手术教育、模拟和数据增强方面具有巨大潜力,然而生成具有真实且临床合理运动的手术视频仍然具有挑战性。大多数现有方法依赖辅助条件(如掩码、轨迹、深度或参考视频)来实现视觉上合理的合成。然而,这些辅助条件通常需要额外的人工标注或专门的采集,使得此类方法难以扩展到小型精选数据集之外。这促使需要一种无参考架构,能够在推理时无需辅助视觉条件即可生成高质量的手术视频。我们提出FLAIR,一种用于无参考手术视频生成的流引导潜动作注入框架。FLAIR从真实手术视频的光流中学习动作先验,根据输入提示动态预测相应的潜动作表示,并将其注入冻结的基础模型中以生成具有改进动作一致性的手术视频。我们进一步构建了SurgActionClip-30K,这是第一个大规模手术视觉数据集,包含以动作中心的分段剪辑和结构化字幕标签,解决了长期缺乏细粒度、以动作中心的手术数据集的问题。最后,我们引入SurgMetrics,这是第一个手术领域专用的评估指标,用于量化生成手术视频的质量,解决了该领域缺乏临床基础评估标准的问题。大量实验表明,FLAIR能够在无需辅助条件的情况下仅通过文本推理生成高质量的手术视频,并且SurgMetrics中的验证表明其与传统指标相比在人类感知一致性方面具有优势。
英文摘要
Surgical video generation holds substantial potential for surgical education, simulation, and data augmentation, yet generating surgical videos with realistic and clinically plausible motion remains challenging. Most existing methods rely on auxiliary conditions, such as masks, trajectories, depth, or reference videos, to achieve visually plausible synthesis. Yet, these auxiliary conditions typically require additional manual annotation or specialized acquisition, making it difficult to scale such methods beyond small, curated datasets. This motivates the need for a reference-free architecture capable of generating high-quality surgical video without requiring auxiliary visual conditions at inference time. We propose FLAIR, a Flow-guided LatentAction Injection framework for Reference-free surgical video generation. FLAIR learns action priors from optical flow of real surgical videos, dynamically predicts corresponding latent action representation from an input prompt, and injects it into a frozen base model to generate surgical videos with improved action consistency. We further construct SurgActionClip-30K, the first large-scale surgical vision dataset comprising action-centric segmented clips and structured caption labels, addressing the persistent lack of fine-grained, action-centric surgical datasets. Lastly, we introduce SurgMetrics, the first surgical domain-specific evaluation metrics for quantifying the quality of generated surgical videos, addressing the persistent absence of clinically grounded evaluation standards in this domain. Extensive experiments demonstrate that FLAIR enables generating high-quality surgical videos using text-only inference without auxiliary conditions, and validation in SurgMetrics demonstrates its strength in alignment with human perception compared to traditional metrics.