arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02471cs.CVcs.AI

基于动作的组织 affordance 支持腹腔镜手术中降低外科医生认知负荷的预期自动构图

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Hongkuan Shi, Qiuyu Yu, Qiang Xie, Hua… 展开作者

Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Hongkuan Shi, Qiuyu Yu, Qiang Xie, Huan Zhao, Wenbin Chen, Caihua Xiong, Chidan Wan, Jing Samantha Pan, Xiong Cai, Han Ding

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出 DiffeoAfford 框架,从手术流程推导视觉注意力监督,训练模型实现 AffordView 自动构图系统,可预测手术区域并降低外科医生认知负荷。

中文摘要 AI 辅助

计算注意力模型可帮助外科医生管理腹腔镜检查的视觉需求,但这类模型需要密集的空间标签,而获取难度大,因为手术意图具有高度专业性和隐含性。本文提出 DiffeoAfford,一种基于动作的组织 affordance 框架,可从已完成的手术流程中回溯推导视觉注意力监督。该框架结合微分同胚约束的组织跟踪与器械轨迹分析,无需逐帧手动标注即可生成 affordance 热点标签。基于这些标签训练的实时预测模型可预测相关手术区域,并实现辅助腹腔镜可视化的自动构图系统 AffordView。该框架与专家标注及术中外科医生视线一致,通过主观、生理和行为测量,在实际评估中降低了外科医生的认知负荷。

英文摘要

In laparoscopy, surgeon gaze tracks where the instruments will act; easing this demand through visual attention modeling requires dense labels of those interaction loci. These encode tacit knowledge: experts converge on consensus loci yet struggle to state the rules. Here we show that such labels can be recovered from completed actions in surgical videos, in which recorded instrument trajectories are converted into dense, continuous supervision. DiffeoAfford grounds tissue affordance by attaching instrument tips to the tissue and transporting them through deformation using diffeomorphism-constrained tracking, matching context-informed annotators' accuracy. Trained on these labels and never on gaze, a real-time model aligns with surgeon gaze more closely in space and time than does camera-assistant gaze. The framework also transfers across procedures: on hysterectomy videos, a separately trained predictor reaches 95.16% directional consistency with subsequent camera motion. In 12 paired cholecystectomies (24 procedures), the auto-framing application AffordView, which proactively centers predicted targets in view, lowered surgeon cognitive workload on converging subjective, physiological, and behavioral measures, including a reduced number of verbal instructions to the camera assistant. Deriving supervision from action rather than manual annotation offers a scalable route to anticipatory assistance.

补充信息

↑