AI 中文总结
该综述梳理2024-2026年手术视频生成领域文献,分三类总结方法,指出生成任务从合成帧转向建模场景因果动态,分析瓶颈并提供实验参考,为相关交叉领域研究者提供参考。
AI 中文摘要
手术视频数据是术中感知、手术流程理解及机器人决策模型的主要训练资源。然而临床数据采集受隐私、成本及类别不平衡限制。手术视频生成已成为解决数据稀缺问题的变革性方法,也是手术模拟、培训及机器人策略学习的基础。该领域发展迅速却缺乏清晰的概念框架。本综述将2024-2026年的文献分为三类:无条件生成、条件生成及世界建模生成,揭示了该任务定义的根本转变——从合成视觉上合理的帧到建模手术场景的因果动态。我们研究了像素级保真度与临床合理性之间持续存在的差距,并确定泛化性、物理真实性、可控性及可解释性为瓶颈。我们还总结了代表性方法在公开数据集上的实验结果,为该领域提供定量参考。本综述对当前状态及开放挑战进行了结构化概述,为从事智能感知、多模态融合、生成式AI及手术数据科学交叉领域研究的人员提供参考。
英文摘要
Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.
Comments4 pages, 1 figures, 3 tables. Accepted for oral presentation at the 2026 3rd International Conference on Intelligent Perception and Pattern Recognition (IPPR 2026)