arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Stitch-Inferencer:通过全景重建增强内窥镜视频分割和跟踪

Stitch-Inferencer: Enhance Endoscopic Video Segmentation and Tracking via Panoramic Reconstruction

Shunsuke Kikuchi, Atsushi Kouno, Hiroki Matsuzaki

arXiv 2607.14968首次发表:更新:

发表机构

Jmees Inc(Jmees公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对内窥镜视频理解中视野受限和遮挡问题,提出Stitch-Inferencer框架,用全景画布取代隐式特征记忆,跨帧拼接观察以扩大视野,让现有模型利用长程上下文,实验证明其能在保持实时性的同时提升分割和跟踪性能。

AI 中文摘要

手术视频理解对导航系统至关重要。内窥镜感知常受视野受限和器械遮挡影响,时空上下文对可靠推理很关键,这促使视频模型跨帧聚合信息。但现有模型常将过去观察隐式存储在学习特征表示中,需特定任务训练、大量标注数据且计算成本高。我们提出Stitch-Inferencer,一个实时、模型无关的推理框架,用显式图像空间全景画布取代隐式特征记忆。通过跨帧拼接有效观察,它在无器械在线视图中保留先前观察像素,扩大有效视野,直接访问当前帧中暂时遮挡或缺失的区域。下游分割或跟踪模型应用于全景图上紧凑的感兴趣区域,预测结果再投影到当前帧,使现有模型无需重新训练就能利用长程上下文。解剖分割和点/框跟踪实验表明,在保持实时吞吐量的同时,能在不同基线水平上持续改进。拼接模块单独运行速度超过60 FPS,为计算受限的术中环境增强内窥镜感知提供了实用的推理时解决方案。源代码将公开。

英文摘要

Surgical video understanding is fundamental to navigation systems. Endoscopic perception is often hindered by a limited field-of-view and frequent instrument occlusions, making spatio-temporal context essential for robust inference. These challenges have motivated video models that aggregate information across frames. However, existing video models typically store past observations implicitly in learned feature representations, often requiring task-specific video training, substantial annotated data, and increased computational cost. We propose Stitch-Inferencer, a real-time, model-agnostic inference framework that replaces implicit feature memory with an explicit image-space panoramic canvas. By stitching valid observations across frames, Stitch-Inferencer preserves previously observed pixels in an online, instrument-free view, expanding the effective field-of-view and providing direct access to regions that are temporarily occluded or absent from the current frame. Downstream segmentation or tracking models are applied to a compact region of interest on the panorama, and their predictions are reprojected to the current frame, enabling existing models to exploit long-range context without retraining. Experiments on anatomy segmentation and point/box tracking demonstrate consistent improvements across diverse baselines while preserving real-time throughput. The stitching module alone runs at over 60 FPS, providing a practical inference-time solution to enhance endoscopic perception in computationally constrained intraoperative environments. Source code will be made publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑