发表机构
University of International Business and Economics; State Key Laboratory of Virtual Reality Technology and Systems; China University of Petroleum; Southwest Jiaotong University; Beijing University Of Technology(对外经济贸易大学; 虚拟现实技术与系统国家重点实验室; 中国石油大学; 西南交通大学; 北京工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对RGB-D SOD中传感器深度不可靠的问题,提出不依赖数据集深度的可靠性感知几何蒸馏框架,利用Depth Anything V2作为教师模型蒸馏几何信息,在多数据集对比中取得最优结果,且几何信息可跨域迁移。
AI 中文摘要
深度信息可解决RGB-D显著目标检测(SOD)中的外观歧义问题,但传感器深度的可靠性并不均匀。缺失区域、模糊边界和结构伪影会通过多模态融合传播,使RGB-D检测器的精度低于仅使用RGB的对应模型。现有的感知质量方法会对观测到的深度进行调控,但仍依赖于同一潜在有缺陷的模态。我们提出了\textit{Method},一种面向RGB-D SOD基准的可靠性感知几何蒸馏框架,在训练和推理期间均不使用数据集提供的深度信息。一个冻结的Depth Anything V2模型仅作为训练时的教师,将密集相对几何、分层空间注意力和边界结构迁移到紧凑的边缘感知几何分支中。池化双向交互将几何与外观对齐,而逐像素可靠性估计器则选择性地注入与当前RGB表示兼容的几何信息。教师模型在训练后被移除,仅保留RGB推理网络。在2985对RGB-掩码上训练后,\textit{Method}在36项指标-数据集对比中取得了26项最优或并列最优的结果,相较于10种最新的RGB-D SOD方法,在ReDWeb-S上实现了13.4%的相对MAE降低。当在DUTS-TR上重新训练时,其在PASCAL-S上将当前最强的F-测度提升了4.2%,表明蒸馏出的几何信息可在特定传感器或数据集域之外迁移。代码将在发表后发布。
英文摘要
Depth can resolve appearance ambiguity in RGB-D salient object detection (SOD), yet sensor depth is not uniformly reliable. Missing regions, blurred boundaries, and structural artifacts can propagate through multimodal fusion and make an RGB-D detector less accurate than its RGB-only counterpart. Existing quality-aware approaches regulate observed depth but remain dependent on the same potentially defective modality. We propose \method, a reliability-aware geometry distillation framework developed for RGB-D SOD benchmarks without using dataset-provided depth during training or inference. A frozen Depth Anything V2 model serves only as a training-time teacher, transferring dense relative geometry, hierarchical spatial attention, and boundary structure to a compact edge-aware geometry branch. Pooled bidirectional interaction aligns geometry with appearance, and a pixel-wise reliability estimator selectively injects geometry that is compatible with the current RGB representation. The teacher is removed after training, leaving an RGB-only inference network. Trained on 2,985 RGB-mask pairs, \method{} achieves the best or tied-best result in 26 of 36 metric-dataset comparisons against ten recent RGB-D SOD methods, including a 13.4\% relative MAE reduction on ReDWeb-S. When retrained on DUTS-TR, it also improves the strongest prior $F$-measure by 4.2\% on PASCAL-S, showing that the distilled geometry transfers beyond a particular sensor or dataset domain. Code will be released upon publication.