手术内窥镜视频的时间视觉语言模型的鲁棒性研究
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
浏览论文内容
中文总结 AI 辅助
该研究针对手术内窥镜视频的时间视觉语言模型,构建Endo-C6基准测试其鲁棒性,提出RobustEndoCLIP,可提升模型在内窥镜损坏下的性能与鲁棒性。
中文摘要 AI 辅助
时间视觉语言模型(TVLMs)为手术视频理解提供了可复用的、基于提示的接口,然而它们在临床现实内窥镜采集伪影下的鲁棒性尚未得到充分表征。实际中,散焦、薄雾、运动模糊、噪声、电凝烟雾和丢包等退化会引入结构化分布偏移,可能损害视频-文本对齐。我们研究时间VLMs在片段帧损坏导致的此类偏移下的鲁棒性,引入Endo-C6——一个包含6种内窥镜真实扰动的紧凑损坏基准,以固定高严重度评估,并将其应用于公开的胃肠道(GI)内窥镜和腹腔镜胆囊切除术视频。在标准化提示协议下,我们对3种近期手术TVLM基线进行基准测试,在均值和最坏情况设置下分析鲁棒性,涵盖294个数据集级评估。最后,我们提出RobustEndoCLIP,通过结合VeRA的少样本参数高效调优得到,其性能优于现有TVLM基线。我们的发现表明,现成TVLMs在内窥镜特定损坏下会出现严重的最坏情况崩溃,而轻量级少样本适配可在不改变基于提示的接口的情况下大幅提升损坏情况下的性能和鲁棒性。我们期望Endo-C6能支持标准化鲁棒性报告,推动更可靠的临床视觉语言系统发展。
英文摘要
Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized. In practice, degradations such as defocus, haze, motion blur, noise, cautery smoke, and packet loss introduce structured distribution shifts which may compromise video-text alignment. We study the robustness of temporal VLMs under such shifts caused by corruptions in clip frames. We introduce Endo-C6, a compact corruption benchmark of six endoscopy-realistic perturbations evaluated at a fixed high severity, and apply it to public Gastrointestinal (GI) endoscopy and laparoscopic cholecystectomy videos. Under a standardized prompt protocol, we benchmark 3 recent surgical TVLM baselines and analyze robustness in both mean and worst-case settings, spanning 294 dataset-level evaluations. Finally, we present RobustEndoCLIP, obtained by few-shot parameter-efficient tuning with VeRA, outperforming existing TVLM baselines. Our findings show that off-the-shelf TVLMs can exhibit severe worst-case collapse under endoscopy-specific corruptions, whereas lightweight few-shot adaptation can substantially improve corrupted performance and robustness without changing the prompt-based interface. We expect Endo-C6 to support standardized robustness reporting and promote more reliable clinical vision-language systems.
发表机构
- Indian Institute of Technology Delhi(印度理工学院德里分校)
- Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
- Birmingham City University(伯明翰城市大学)
- University Hospitals Birmingham(伯明翰大学医院)
- Khalifa University(哈利法大学)
- German Cancer Research Center (DKFZ)(德国癌症研究中心)
机构由 AI 辅助整理,请以论文原文为准。