arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19889cs.CV

LAVIFT:用于手术交互识别的潜在动作引导视觉微调

Latent-Action-Guided Video-Language Feature Learning for Surgical Instrument-Tissue Interaction Recognition

  • Arizona State University(亚利桑那州立大学)
  • University of California, Merced(加州大学默塞德分校)
  • Intel Corporation(英特尔公司)

机构由 AI 辅助整理,请以论文原文为准。

Jiajun Cheng, Sainan Liu, Subarna Tripathi, Xiaofan Yu, Shan Lin

AI总结:

研究针对手术交互识别中预训练模型适应细粒度交互的挑战,提出LAViFiT框架,通过逆动力学模型、前向世界模型及补丁级SIG正则化器,实现视觉语言微调,提升了识别率和图像-文本对齐,增强了特征基础和空间连贯性。

AI中文摘要:

理解器械与组织的交互对于情境感知手术人工智能和自主机器人手术至关重要。预训练的视觉语言模型(VLM)和视觉编码器通过转移广泛的视觉和语义知识,为传统交互分类器提供了替代方案。然而,使其适应细粒度手术交互仍具有挑战性:冻结视觉编码器完全依赖可能保留噪声且空间定位较弱的预训练表示;完全微调可改善全局语义对齐,但不能确保编码器在正确动作区域学习有意义的特征。我们引入LAViFiT来解决这些限制,它是一个用于视觉语言微调的端到端潜在动作引导框架。逆动力学模型捕获每个动作引起的视觉变化,前向世界模型驱动编码器表示与动作相关的区域。补丁级SIG正则化器在无额外监督(如边界框或伪标签)的情况下进一步防止局部特征崩溃。跨多个编码器和数据集的实验提高了识别率和图像-文本对齐,同时表示分析显示在完整的器械-组织交互区域上有更强的基础和更空间连贯的特征。

英文摘要:

Recognizing instrument--tissue interactions is essential for context-aware surgical AI. Vision-language models offer a natural way to inject semantic structure into surgical representations by aligning video features with textual action descriptions. However, pretrained encoders may lack spatial coherence, while global semantic alignment does not ensure precise spatial and temporal representations. By analyzing frame-to-frame feature changes, we find that semantic alignment increases their dimensionality, but larger increases do not necessarily improve recognition; encoders also differ in how strongly dominant changes localize to interaction regions. Motivated by these findings, we introduce \ours{}, which compresses frame-to-frame changes into latent actions and predicts next-frame features during end-to-end video--language alignment. Without additional spatial or motion annotations, \ours{} improves the interaction grounding of leading feature changes and temporal-direction sensitivity in our evaluated settings. We further characterize how action capacity and prediction strength affect recognition across encoders and triplet components. Using image encoders without large-scale video pretraining, \ours{} achieves competitive recognition with faster inference and smaller INT4 accuracy drops than V-JEPA2/2.1.

↑