arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当深度信息产生负面影响:面向无深度RGB-D显著目标检测的可靠性感知几何蒸馏

When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection

Xuehao Wang, Jiaxin Hua, Runmei Li, Zhenyu Wu, Chenglizhao Chen, Ke Gu, Aimin Hao

arXiv 2609.03378首次发表:更新:

发表机构

University of International Business and Economics; State Key Laboratory of Virtual Reality Technology and Systems; China University of Petroleum; Southwest Jiaotong University; Beijing University Of Technology(对外经济贸易大学; 虚拟现实技术与系统国家重点实验室; 中国石油大学; 西南交通大学; 北京工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对RGB-D SOD中传感器深度不可靠的问题,提出不依赖数据集深度的可靠性感知几何蒸馏框架,利用Depth Anything V2作为教师模型蒸馏几何信息,在多数据集对比中取得最优结果,且几何信息可跨域迁移。

AI 中文摘要

深度信息可解决RGB-D显著目标检测(SOD)中的外观歧义问题,但传感器深度的可靠性并不均匀。缺失区域、模糊边界和结构伪影会通过多模态融合传播,使RGB-D检测器的精度低于仅使用RGB的对应模型。现有的感知质量方法会对观测到的深度进行调控,但仍依赖于同一潜在有缺陷的模态。我们提出了\textit{Method},一种面向RGB-D SOD基准的可靠性感知几何蒸馏框架,在训练和推理期间均不使用数据集提供的深度信息。一个冻结的Depth Anything V2模型仅作为训练时的教师,将密集相对几何、分层空间注意力和边界结构迁移到紧凑的边缘感知几何分支中。池化双向交互将几何与外观对齐,而逐像素可靠性估计器则选择性地注入与当前RGB表示兼容的几何信息。教师模型在训练后被移除,仅保留RGB推理网络。在2985对RGB-掩码上训练后,\textit{Method}在36项指标-数据集对比中取得了26项最优或并列最优的结果,相较于10种最新的RGB-D SOD方法,在ReDWeb-S上实现了13.4%的相对MAE降低。当在DUTS-TR上重新训练时,其在PASCAL-S上将当前最强的F-测度提升了4.2%,表明蒸馏出的几何信息可在特定传感器或数据集域之外迁移。代码将在发表后发布。

英文摘要

Depth can resolve appearance ambiguity in RGB-D salient object detection (SOD), yet sensor depth is not uniformly reliable. Missing regions, blurred boundaries, and structural artifacts can propagate through multimodal fusion and make an RGB-D detector less accurate than its RGB-only counterpart. Existing quality-aware approaches regulate observed depth but remain dependent on the same potentially defective modality. We propose \method, a reliability-aware geometry distillation framework developed for RGB-D SOD benchmarks without using dataset-provided depth during training or inference. A frozen Depth Anything V2 model serves only as a training-time teacher, transferring dense relative geometry, hierarchical spatial attention, and boundary structure to a compact edge-aware geometry branch. Pooled bidirectional interaction aligns geometry with appearance, and a pixel-wise reliability estimator selectively injects geometry that is compatible with the current RGB representation. The teacher is removed after training, leaving an RGB-only inference network. Trained on 2,985 RGB-mask pairs, \method{} achieves the best or tied-best result in 26 of 36 metric-dataset comparisons against ten recent RGB-D SOD methods, including a 13.4\% relative MAE reduction on ReDWeb-S. When retrained on DUTS-TR, it also improves the strongest prior $F$-measure by 4.2\% on PASCAL-S, showing that the distilled geometry transfers beyond a particular sensor or dataset domain. Code will be released upon publication.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑