arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩展KAFR:用于手术视频高效分析的运动学自适应范式

Extended KAFR: A kinematic-adaptive paradigm for the efficient analysis of surgical video

Huu Phong Nguyen, Shekhar Madhav Khairnar, Ganesh Sankaranarayanan

arXiv 2608.01058首次发表:更新:

发表机构

University of Texas Southwestern Medical Center(德克萨斯大学西南医学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究验证了运动学自适应帧识别(KAFR)可泛化至腹腔镜手术,其结合YOLO、X3D模型,仅用0.58%的帧就达到91.0%的F1分数,性能优于传统采样且与先进模型相当,大幅降低了手术视频分析的计算负担。

AI 中文摘要

人工智能正越来越多地应用于手术视频分析,以实现阶段分割、技能评估和工作流程优化。一个关键挑战是手术录像的长度通常为1到数小时,这会产生巨大的计算负担。我们此前开发了用于机器人手术的运动学自适应帧识别(KAFR),该方法表明,跟踪器械运动可有效识别信息帧,同时过滤冗余内容。然而,腹腔镜手术带来了额外挑战:手动控制摄像头会导致频繁的运动伪影,且图像质量通常低于机器人系统。本研究使用包含80例腹腔镜胆囊切除术(标注了7个手术阶段)的Cholec80基准,评估KAFR是否可泛化至腹腔镜手术。KAFR分为三个阶段:微调后的YOLO模型检测并分割手术器械;基于器械位移或速度变化自适应选择帧;X3D模型将选定帧分类为手术阶段。KAFR仅使用0.58%的帧进行阶段分类,F1分数达到91.0%,与典型的4%帧采样相比减少了约7倍,同时保持了与LoViT(90.2%)和Trans-SVNet(89.7%)相当的性能。这些结果表明,基于运动学的帧选择可有效迁移至具有挑战性的腹腔镜环境。

英文摘要

Artificial Intelligence is increasingly applied to surgical video analysis for phase segmentation, skill assessment, and workflow optimization. A key challenge is the length of surgical recordings, often one to several hours, creating substantial computational burden. We previously developed Kinematics-Adaptive Frame Recognition (KAFR) for robotic surgery, showing that tracking tool motion effectively identifies informative frames while filtering redundant content. However, laparoscopic surgery introduces additional challenges: manual camera control causes frequent motion artifacts, and image quality is generally lower than robotic systems. This study evaluates whether KAFR generalizes to laparoscopic surgery using the Cholec80 benchmark, comprising 80 laparoscopic cholecystectomy procedures annotated for seven surgical phases. KAFR operates in three stages: a fine-tuned YOLO model detects and segments surgical tools; frames are adaptively selected based on tool displacement or velocity variation; and an X3D model classifies selected frames into surgical phases. KAFR achieved a 91.0\% F1 score using only 0.58\% of frames for phase classification, representing an approximately seven-fold reduction compared to typical 4\% frame sampling, while maintaining performance comparable to LoViT (90.2\%) and Trans-SVNet (89.7\%). These results demonstrate that kinematics-based frame selection transfers effectively to the challenging laparoscopic environment.

Comments18

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑