发表机构
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对单目深度估计的视觉歧义问题,提出CapDepth框架,通过详细长文本引导,在非朗伯表面和恶劣天气下的深度误差较现有最优方法分别降低25.0%和22.0%。
AI 中文摘要
单目深度估计(Monocular Depth Estimation, MDE)因单张图像有限信息固有的视觉歧义,面临非朗伯表面和恶劣天气条件的挑战。现有工作通过图像修复或增强单独解决这些问题,带来的鲁棒性提升有限。语言作为视觉的强大互补模态,已被证明可通过详细长文本增强视觉-语言模型(Vision-Language Models, VLMs)的视觉感知能力。然而,现有结合语言的MDE方法未能充分利用这一潜力,原因在于文本输入短、信息有限、全局文本特征学习粗糙,以及深度解码过程中语言引导不足。为解决这些局限,我们提出CapDepth,一种用于鲁棒MDE的新型框架,利用详细长文本的引导缓解挑战性场景中的视觉歧义。首先,我们设计详细长文本输入模板,明确传递多个原子句间的丰富空间关系;其次,引入动态文本编码器,通过渐进式掩码注意力提取与深度相关的细粒度文本特征;最后,提出文本自适应解码器,通过稳定自适应层归一化,利用文本特征引导增强的深度解码。大量实验验证了CapDepth的有效性,其性能优于现有最优方法,在非朗伯表面上深度误差降低25.0%,在恶劣天气条件下深度误差降低22.0%。
英文摘要
Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent in single-image limited information. Existing works address them in isolation via image inpainting or augmentation, yielding limited robustness gains. Language, as a powerful complementary modality to vision, is demonstrated to enhance the visual perception capabilities of vision-language models (VLMs) via detailed long captions. However, prior language-integrated MDE methods fail to fully harness this potential due to short text input with limited information, coarse global text feature learning, and limited language guidance during depth decoding. To address these limitations, we propose CapDepth, a novel framework for robust MDE that leverages guidance from detailed long captions to alleviate visual ambiguities in both challenging scenarios. First, we design a detailed long caption input template that explicitly conveys rich spatial relationships among multiple atom sentences. Second, a dynamic caption encoder is introduced to extract fine-grained depth-relevant text features via progressive masked attention. Finally, we propose a text-adaptive decoder that guides enhanced depth decoding with text features via stable adaptive layer normalization. Extensive experiments validate the efficacy of CapDepth, which outperforms state-of-the-art methods, achieving depth error reductions of 25.0% on non-Lambertian surfaces and 22.0% under adverse weather conditions.
CommentsAccepted to ACM MM 2026