arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RGBD20K:用于RGB-D语义分割的大规模基准数据集

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan

arXiv 2609.29028首次发表:更新:

发表机构

University of North Texas(北德克萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出RGBD20K大规模RGB-D语义分割基准,含160类、2万对图像及高保真标注,并引入SPF融合方法,在多个基准上达到最先进性能。

AI 中文摘要

在本文中,我们提出了RGBD20K,这是一个新颖的数据集,通过涵盖丰富的类别和高精度的标注,促进更鲁棒和通用的RGB-D语义分割的发展。RGBD20K具有几个吸引人的特性:(1)扩展的语义空间。具体而言,它涵盖了160个细粒度类别,大大超过了现有流行RGB-D基准(例如,具有40个类别的NYUv2和具有37个类别的SUN RGB-D)的类别多样性。凭借如此丰富的语义覆盖,我们期望促进更可泛化的分割模型的学习。(2)更大规模。与当前基准相比,RGBD20K提供了20,000对RGB-D图像,提供了更大的训练资源,有利于开发更强大的深度模型。(3)高保真标注。我们对现有标签进行了严格的重新评估和修正,以解决长期存在的标注噪声,从而形成干净可靠的真实基础。此外,我们提出了一种新颖的分数净化融合(SPF)方法,在所有评估基准上均取得了最先进的性能,证明了我们方法在利用高质量多模态信息进行RGB-D语义分割方面的有效性。数据集位于:此https URL。

英文摘要

In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-grained categories, largely surpassing the category diversity of existing popular RGB-D benchmarks (e.g., NYUv2 with 40 classes and SUN RGB-D with 37 classes). With such enriched semantic coverage, we expect to promote the learning of more generalizable segmentation models. (2) Larger Scale. Compared with current benchmarks, RGBD20K offers 20,000 RGB-D image pairs, providing a substantially larger training resource that benefits the development of more powerful deep models. (3) High-Fidelity Annotation. We perform rigorous re-evaluation and correction of existing labels to resolve long-standing annotation noise, resulting in a clean and reliable ground-truth foundation. Furthermore, we propose a novel score-purified fusion (SPF) method, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of our approach in leveraging high-quality multimodal information for RGB-D semantic segmentation. The dataset is here: https://github.com/ShaohuaDong2021/RGBD20K/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑