arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RDVSv2:用于RGB-D视频显著目标检测的大规模基准测试

RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

Tianyu Li, Jiahao He, Keren Fu, Qijun Zhao

arXiv 2607.25392首次发表:更新:

发表机构

National Key Laboratory of Fundamental Science on Synthetic Vision, Sichuan University; College of Computer Science, Sichuan University(四川大学合成视觉基础科学国家重点实验室; 四川大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍用于RGB-D视频显著目标检测的大规模基准测试RDVSv2,其规模大、场景多样。基于SAM2建立强大基线,采用参数高效微调策略联合编码多模态线索,该基线在RDVSv2及现有基准测试中达最优,为相关研究提供资源。

AI 中文摘要

我们引入了RDVSv2,这是一个用于RGB-D视频显著目标检测(RGB-D VSOD)的大规模基准测试,带有密集的帧级注释。该领域现有数据集在规模和注释质量上往往有限,且依赖几何一致性较差的深度线索。为解决这些局限,RDVSv2基于可公开获取的立体在线视频构建,包含249个视频序列及29077个注释帧,有来自立体视频的深度图和基于眼动追踪引导注释的逐帧显著目标掩码。与现有数据集相比,它规模更大且场景更多样、更具挑战性。此外,我们基于Segment Anything Model 2(SAM2)为RGB-D VSOD建立了一个强大的基线,采用参数高效微调策略使SAM2编码器联合编码RGB、深度和光流线索。大量实验表明,RDVSv2对现有方法更具挑战性,所提基线在RDVSv2和现有基准测试中取得了最优结果。我们希望RDVSv2及提供的基线能为未来RGB-D VSOD及相关多模态视频理解任务的研究提供有用资源。我们的数据集和代码将在指定网址提供。

英文摘要

We introduce RDVSv2, a large-scale benchmark for RGB-D video salient object detection (RGB-D VSOD) with dense frame-level annotations. Existing datasets in this emerging field are often limited in scale and annotation quality, while also relying on less geometry-consistent depth cues. To address these limitations, RDVSv2 is built from publicly accessible stereoscopic online videos and contains 249 video sequences with 29,077 annotated frames. It includes depth maps derived from stereoscopic videos, together with frame-wise salient object masks annotated with eye-tracking guidance. Compared with existing datasets, RDVSv2 is much larger in scale and covers more diverse and challenging scenarios. In addition, we establish a strong baseline for RGB-D VSOD based on Segment Anything Model 2 (SAM2). Specifically, we employ a parameter-efficient fine-tuning (PEFT) strategy to adapt the SAM2 encoder to jointly encode RGB, depth, and optical flow cues. Extensive experiments show that RDVSv2 is substantially more challenging for existing RGB-D VSOD methods. Meanwhile, the proposed baseline achieves state-of-the-art results on RDVSv2 and existing RGB-D VSOD benchmarks. We hope that RDVSv2 and the provided baseline will serve as useful resources for future research on RGB-D VSOD and related multi-modal video understanding tasks. Our dataset and code will be available at https://github.com/ltynick/RDVSv2.

CommentsAccepted to ACMMM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑