arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向自然场景的深度主导型骨架检测

Depth-Dominant Skeleton Detection for Natural Scenes

Chengkun Rao, Yixuan Deng, Min Li, Yangjun Ou, Ye Li, Ziwei Luo, Zhaojing Wang, Junwei Tang, Bangchao Wang, Xiaoyun Yan

arXiv 2608.16367首次发表:更新:

发表机构

School of Computer Science and Artificial Intelligence, Wuhan Textile University(武汉纺织大学计算机科学与人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有自然场景骨架检测方法在复杂图像上性能下降的问题,提出以深度为主、RGB为辅的新范式,构建轻量模型DDSkel,在SymPASCAL数据集上以更少参数实现了优于SOTA的性能。

AI 中文摘要

迄今为止,所有自然场景骨架检测均遵循以RGB图像作为唯一输入的范式;尽管已取得显著进展,但该范式下的方法在内容复杂的图像上性能会大幅下降。我们观察到,深度图像对颜色和纹理天生不敏感,能提供清晰的区域轮廓及区域间空间关系,自然缓解了复杂场景下骨架检测的难度。受此观察启发,本文首次提出一种新颖的骨架检测范式:以深度图像作为主导模态,RGB图像作为辅助模态;并据此提出该范式下的模型DDSkel(Depth-Dominant Skeleton Detection的缩写)。DDSkel采用不对称编码器设计,将RGB信息融合到深度特征中,其中RGB模态分支的参数仅为深度模态分支的12%。DDSkel结构简单,无复杂设计;然而,仅使用当前最优方法36%的可训练参数,便在最具挑战性、包含大量复杂图像的数据集SymPASCAL上,性能优于所有现有最优方法。

英文摘要

To date, all natural scene skeleton detection follows the paradigm of taking RGB images as the sole input; despite notable progress, methods under this paradigm suffer significant performance degradation on complex-content images. We observe that depth images are inherently insensitive to color and texture, and can provide clear regional contours and inter-region spatial relationships, which naturally alleviates the difficulty of skeleton detection in complex scenarios. Motivated by this observation, this paper proposes for the first time a novel skeleton detection paradigm where depth images serve as the dominant modality and RGB images act as the auxiliary, and accordingly presents a model DDSkel (short for Depth-Dominant Skeleton Detection) under this paradigm. DDSkel employs an asymmetric encoder design to fuse RGB information into depth features, with the RGB modality branch having only 12% the parameters of the depth modality branch. DDSkel has a simple structure without intricate designs. Nevertheless, with only 36% of the trainable parameters of the current best method, DDSkel outperforms all state-of-the-art approaches on SymPASCAL, the most challenging dataset with a large volume of complex images.

Comments11 pages, 3 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑