AI 中文总结
针对母子ERCP中双柔性镜无校准立体关系的手术视频预测问题,提出角色非对称的CrossScope双流模型,经配对双视野基准评估,其性能优于现有基线。
AI 中文摘要
视觉世界模型通常从单一观测流学习未来动态,这限制了它们对具有多个独立移动观测者的协作系统的建模能力。我们在母子内窥镜逆行胰胆管造影(ERCP)中研究这一挑战,其中两个柔性镜提供互补但依赖角色的视野,且无校准的立体关系。与假设对称信息交换的传统多视图融合不同,我们提出角色非对称双视野未来预测,即根据预测目标及其潜在空间需求选择性传递跨视图证据。我们提出CrossScope,这是一种双流手术世界模型,保留视图特定专家,同时通过几何引导的残差交互实现目标特定证据路由。CrossScope学习两种互补的通信方向:来自母镜的几何运动线索引导子镜的未来动态;仅当建立有效空间对应关系时,姿态对齐的子镜外观才支持母镜预测。该设计允许每个镜贡献任务相关证据,同时不损害其视图特定表示。为评估该问题,我们建立了配对双视野基准,包含同步的体模和真实世界ERCP片段,评估指标包括视觉保真度、结构保留、目标定位和运动一致性。实验表明,CrossScope始终优于强大的手术视频生成基线,验证了角色感知证据路由对多观测者视觉世界建模的重要性。
英文摘要
Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother--Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship. Unlike conventional multi-view fusion that assumes symmetric information exchange, we formulate \textbf{role-asymmetric dual-scope future prediction}, where cross-view evidence is selectively transferred according to the prediction target and its underlying spatial requirements. We propose \textbf{CrossScope}, a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions. CrossScope learns two complementary communication directions: geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. This design allows each scope to contribute task-relevant evidence without compromising its view-specific representation. To evaluate this problem, we establish a paired dual-scope benchmark comprising synchronized phantom and real-world ERCP episodes, with evaluations assessing visual fidelity, structural preservation, target localization, and motion consistency. Experiments demonstrate that CrossScope consistently outperforms strong surgical video generation baselines, validating the importance of role-aware evidence routing for multi-observer visual world modeling.