可聚焦的单目深度估计
Focusable Monocular Depth Estimation
- School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出FDE,一种区域感知的单目深度估计方法,通过FocusDepth框架和MSSA模块,提升目标区域的深度精度和全局几何一致性,基于FDE-Bench验证了方法的有效性。
AI中文摘要:
单目深度基础模型在不同场景中表现良好,但通常使用统一的像素级优化目标,无法区分用户指定或任务相关的目标区域与周围环境。为此,我们引入可聚焦的单目深度估计(FDE),一种区域感知的深度估计任务,要求在给定目标区域时,模型优先保证前景深度精度,保持清晰的边界过渡,并维持一致的全局场景几何。为优先处理任务关键区域建模,我们提出了FocusDepth,一种基于提示的单目相对深度估计框架,通过框/文本提示引导深度模型聚焦于目标区域。FocusDepth的核心多尺度空间对齐融合(MSSA)将Segment Anything Model 3的多尺度特征空间对齐到Depth Anything家族,并通过尺度特定的门控条件融合注入其中。这使得能够实现密集的提示信息注入而不破坏几何表示,从而赋予深度估计模型聚焦感知能力。为研究FDE,我们建立了FDE-Bench,一个以目标为中心的单目相对深度基准,来源于五个数据集中的图像-目标-深度三元组,包含252.9K/72.5K训练/验证三元组和972个类别,涵盖现实世界和体素模拟环境。在FDE-Bench上,FocusDepth在框和文本提示下均优于全局微调的DA2/DA3基线,最大增益出现在目标边界和前景区域,同时保持全局场景几何。消融实验显示,MSSA的空间对齐是关键设计因素,破坏提示-几何对应关系会使AbsRel增加高达13.8%。
英文摘要:
Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-relevant target regions from the surrounding context. We therefore introduce Focusable Monocular Depth Estimation (FDE), a region-aware depth estimation task in which, given a specified target region, the model is required to prioritize foreground depth accuracy, preserve sharp boundary transitions, and maintain coherent global scene geometry. To prioritize task-critical region modeling, we propose FocusDepth, a prompt-conditioned monocular relative depth estimation framework that guides depth modeling to focus on target regions via box/text prompts. The core Multi-Scale Spatial-Aligned Fusion (MSSA) in FocusDepth spatially aligns multi-scale features from Segment Anything Model 3 to the Depth Anything family and injects them through scale-specific, gated conditional fusion. This enables dense prompt cue injection without disrupting geometric representations, thereby endowing the depth estimation model with focused perception capability. To study FDE, we establish FDE-Bench, a target-centric monocular relative depth benchmark built from image-target-depth triplets across five datasets, containing 252.9K/72.5K train/val triplets and 972 categories spanning real-world and embodied simulation environments. On FDE-Bench, FocusDepth consistently improves over globally fine-tuned DA2/DA3 baselines under both box and text prompts, with the largest gains appearing in target boundary and foreground regions while preserving global scene geometry. Ablations show that MSSA's spatial alignment is the key design factor, as disrupting prompt-geometry correspondence increases AbsRel by up to 13.8%.