发表机构
The Chinese University of Hong Kong; Shenzhen Loop Area Institute; the University of Sydney(香港中文大学; 深圳河套学院; 悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Surg-UniWorld统一外科世界模型,构建Chlec80-SurgWAM基准,经实验证实其在可控外科视频生成相关指标上优于现有方法。
AI 中文摘要
可控外科世界模型可通过合成真实的器械-组织交互,为外科人工智能与模拟提供生成基础。然而现有方法缺乏统一的多模态控制范式,而异质视觉条件的直接融合常导致解剖结构失真、器械外观漂移及时间上不一致的交互。本研究提出Surg-UniWorld,一种带有多模态控制专家的统一外科世界模型。Surg-UniWorld首先从首帧外观与分层语义掩码构建分层外科锚点,以保留持久场景身份、解剖结构及交互边界;随后,锚点相关多模态专家解释相对于共享锚点的边缘、深度及光流证据,捕捉互补的边界、几何与运动信息;多模态控制专家进一步对激活的模态增量进行贡献保留的阶段性组合,并为Wan2.2视频扩散主干生成控制提示。为支持多模态外科世界建模,本研究还构建了用于可控外科视频生成的基准Chlec80-SurgWAM。大量实验表明,Surg-UniWorld在生成质量、时间一致性及多模态可控性上均持续优于现有可控视频生成方法与外科世界模型基线。
英文摘要
Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions. However, existing methods lack a unified multimodal control paradigm, while direct fusion of heterogeneous visual conditions often causes anatomical distortion, instrument appearance drift, and temporally inconsistent interactions. In this work, we propose {Surg-UniWorld}, a unified surgical world model with multimodal control experts. Surg-UniWorld first constructs a {Hierarchical Surgical Anchor} from first-frame appearance and hierarchical semantic masks to preserve persistent scene identity, anatomical organization, and interaction boundaries. {Anchor-Relative Modality Experts} then interpret edge, depth, and optical-flow evidence relative to the shared anchor, capturing complementary boundary, geometric, and motion information. A {Multimodal Control Expert} further performs contribution-preserving stage-wise composition of the activated modality increments and generates control hints for the Wan2.2 video diffusion backbone. To support multimodal surgical world modeling, we further construct Cholec80-SurgWAM, a benchmark for controllable surgical video generation. Extensive experiments demonstrate that Surg-UniWorld consistently outperforms existing controllable video generation methods and surgical world-model baselines in generation quality, temporal consistency, and multimodal controllability.