发表机构
Hefei University of Technology; Nanjing University of Posts and Telecommunications; Chengpin Home Tech; Differential Robotics(合肥工业大学; 南京邮电大学; 橙品家居科技; 差分机器人)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出FlexDepth,一种尺度驱动的自监督单目深度估计模型系列,通过两阶段静态-动态解耦训练和尺度驱动解码器,在复杂驾驶场景中实现高精度深度估计,且计算开销低。
AI 中文摘要
自监督单目深度估计(MDE)近年来因其不依赖真实深度而受到关注。然而,大多数现有模型局限于单一尺度,在复杂驾驶环境中性能显著下降。专门处理动态交通参与者的网络往往过于复杂,阻碍其在资源受限的汽车边缘设备上部署。为解决这些限制并迈向鲁棒的驾驶感知,我们提出了FlexDepth,一个尺度驱动的、灵活的自监督MDE模型系列,专为具有挑战性的道路场景设计。FlexDepth采用两阶段静态-动态解耦训练策略,能够独立评估静态背景和动态道路目标的置信度。此外,它引入了一个精心设计的尺度驱动解码器(SDD),根据尺度大小动态选择组件,促进高效特征融合并输出高精度深度图。在标准驾驶基准上的大量实验表明,无需任何辅助信息,我们的模型在任意尺度下均以最小的计算开销实现了最先进的性能。我们最小的模型Flex-Nano仅需0.7 GFLOPs,在移动平台上达到37.6 FPS,在保持优秀零样本泛化能力的同时确保可靠的实时感知。源代码可在以下网址获取:this https URL
英文摘要
Self-Supervised Monocular Depth Estimation (MDE) has garnered attention in recent years due to its independence from ground truth. However, most existing models are limited to a single scale and exhibit considerable performance degradation in complex driving environments. Networks specifically designed to handle dynamic traffic participants tend to be overly complex, hindering their deployment on resource-constrained automotive edge devices. To address these limitations and move towards robust driving perception, we propose FlexDepth, a scale-driven and flexible family of self-supervised MDE models tailored for challenging road scenarios. FlexDepth employs a two-stage static-dynamic decoupled training strategy, enabling the independent assessment of confidence for both static backgrounds and dynamic road objects. Furthermore, it introduces a meticulously designed Scale-Driven Decoder (SDD) to dynamically select components based on scale size, facilitating efficient feature fusion and the output of high-precision depth maps. Extensive experiments on standard driving benchmarks demonstrate that without any auxiliary information, our model achieves state-of-the-art performance across arbitrary scales with minimal computational overhead. Our smallest model, Flex-Nano, requires only 0.7 GFLOPs and achieves 37.6 FPS on mobile platforms, ensuring reliable real-time perception while maintaining excellent zero-shot generalization. Our source code is available: https://github.com/startnew/flexdepth
CommentsAccepted by ECCV2026. Code is available at https://github.com/startnew/flexdepth
Journal refComputer Vision - ECCV 2026, Lecture Notes in Computer Science, vol. 17003, Springer, Cham, 2026
DOI:10.1007/978-3-032-36846-1_7