发表机构
University of Thessaly(塞萨利大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出自监督可感知解释的单目深度估计框架PIMDE,通过分解图像为感知特征图并独立估计后融合,在KITTI上实现与现有方法相当的精度,同时增强可解释性。
AI 中文摘要
自监督单目深度估计(MDE)能够从单目图像中预测深度,无需真实标签监督,因此在大规模及现实世界应用中具有吸引力。尽管精度持续提升,但大多数现有方法仍难以解释,因为深度是从RGB表示中推断出来的,这些表示掩盖了单个感知图像成分的影响。这种缺乏透明度限制了系统性地分析失败案例,并降低了安全关键场景中的置信度。本文提出了一种用于可感知解释的单目深度估计(PIMDE)的自监督框架,旨在将深度预测与输入图像的各个感知成分相关联。所提方法不是直接处理RGB输入,而是将每幅图像分解为一组感知特征图(PFM),每个特征图编码一个特定的视觉线索。不同的深度估计分支独立处理这些PFM以产生深度估计(PIDE),随后通过显式融合策略进行组合。这种表述使我们能够直接检查每个感知线索对最终深度预测的贡献。在KITTI基准数据集上进行的实验表明,PIMDE在性能上与已建立的自监督MDE方法相当,同时提供了关于不同感知线索如何影响深度估计的额外见解。这些结果表明,感知分解可以在不牺牲深度估计精度的情况下支持可解释性。
英文摘要
Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, as depth is inferred from RGB representations that obscure the impact of individual perceptual image components. This lack of transparency limits systematic analysis of failure cases and reduces confidence in safety-critical settings. This paper presents a self-supervised framework for perceptually interpretable monocular depth estimation (PIMDE), designed to associate depth predictions with distinct perceptual components of the input image. Rather than operating directly on RGB inputs, the proposed method decomposes each image into a set of perceptual feature maps (PFMs), each encoding a specific visual cue. Distinct depth estimation branches process these PFMs independently to produce depth estimates (PIDEs), which are subsequently combined through an explicit fusion strategy. This formulation allows us to examine directly the contribution of each perceptual cue to the final depth prediction. Experiments conducted on the KITTI benchmark dataset demonstrate that PIMDE achieves performance comparable to established self-supervised MDE methods while providing additional insight into how different perceptual cues influence depth estimation. These results indicate that perceptual decomposition can support interpretability without sacrificing depth estimation accuracy.
CommentsPublished at IEEE ICIP 2026; 6 pages, 4 figures
Journal ref2026 IEEE International Conference on Image Processing (ICIP), pp. 1-6, 2026
DOI:10.1109/ICIP61757.2026.11630094