LD-RSVIS:用于指代手术视频器械分割的大规模多样化基准
LD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument Segmentation
浏览论文内容
中文总结 AI 辅助
针对现有指代手术视频器械分割基准规模小且仅支持单目标的问题,提出大规模多样化基准LD-RSVIS,含3536个视频、30类器械及多目标/无目标设置,并评估12种方法,提出Cascade-RSVIS方法。
中文摘要 AI 辅助
指代手术视频器械分割(RSVIS)旨在根据文本描述对手术视频中的器械进行分割。尽管近期取得了进展,但当前模型是在相对小规模的基准上进行训练和评估的,这阻碍了更通用的RSVIS的发展。此外,现有基准仅支持指代视频中单个器械的单目标表达,而忽略了多目标和无目标指代表达,限制了RSVIS在实际场景中的适用性。针对这些问题,我们提出了LD-RSVIS,一个新基准,旨在促进更稳健和更通用的RSVIS。具体而言,LD-RSVIS包含3,536个手术视频,共109万帧,覆盖来自25种不同手术过程的30个器械类别。通过包含丰富的视频和类别,LD-RSVIS有利于大规模训练和评估更通用的RSVIS方法。此外,与现有数据集不同,LD-RSVIS提供多样化的指代设置,包括无目标、单目标和多目标表达,这使得开发更适用于实际应用的RSVIS模型成为可能。为确保高质量标注,LD-RSVIS中的所有视频均经过人工标注,并进行了多轮检查和细化。据我们所知,LD-RSVIS是规模最大、最多样化的RSVIS基准。为了分析LD-RSVIS并为未来研究提供比较,我们评估了12种代表性方法,结果表明仍需更多努力来改进。为鼓励未来研究,我们提出了一种简单而有效的RSVIS方法,称为Cascade-RSVIS,该方法首先利用互补的多线索文本信息挖掘目标特定线索,然后利用这些线索和文本信息进行分割,取得了有前景的性能。我们的基准和代码将公开发布。
英文摘要
Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models are trained and assessed on relatively small-scale benchmarks, hindering the development of more general RSVIS. In addition, existing benchmarks support only the single-target expression that refers to one instrument in the video, while overlooking multi-target and no-target referring expressions, restricting the applicability of RSVIS in practical scenarios. Addressing these issues, we propose LD-RSVIS, a new benchmark aiming to facilitate more robust and general RSVIS. Specifically, LD-RSVIS consists of 3,536 surgical videos with 1.09 million frames and covers a broad set of 30 instrument classes from 25 various procedures. By including abundant videos and classes, LD-RSVIS could benefit large-scale training and evaluation of more general RSVIS methods. Besides, unlike existing datasets, LD-RSVIS offers diverse referring settings, including no-target, single-target, and multi-target expressions, which enables the development of more practical RSVIS models in real applications. In order to ensure high-quality annotations, all videos in LD-RSVIS are manually labeled with multiple rounds of inspection and refinement. To our knowledge, LD-RSVIS is the largest and most diverse benchmark for RSVIS. To analyze LD-RSVIS and to provide comparison for future research, we evaluate 12 representative methods, and the results reveal that more efforts are required for improvements. To encourage future research, we present a simple yet effective RSVIS method, dubbed Cascade-RSVIS, that first mines target-specific cues using the complementary multi-cue text information and then employs such cues and textual information for segmentation, achieving promising performance. Our benchmark and code will be released.
发表机构
- University of North Texas(北德克萨斯大学)
- Meta Inc(Meta公司)
机构由 AI 辅助整理,请以论文原文为准。