PXDepth:用于结构保留单目深度估计的像素空间建模
PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation
浏览论文内容
中文总结 AI 辅助
该研究针对单目深度估计难以保留细结构的问题,提出PXDepth模型,通过分离全局上下文建模与像素级预测,在零样本基准上实现了兼具局部几何精度与全局深度一致性的高效估计。
中文摘要 AI 辅助
现有单目深度估计器已实现出色的零样本泛化能力,但常难以保留细粒度结构与物体边界。我们将此局限归因于大补丁ViT编码器与卷积解码器的普遍组合,因为粗粒度分词会削弱上采样无法完全恢复的像素级线索。为解决该问题,我们提出PXDepth,一种判别式单目深度模型,其将全局上下文建模与像素级深度预测分离。具体而言,大补丁ViT捕获全局场景上下文,而由上下文调制像素Transformer块构成的像素空间预测器在整个深度估计过程中维持高分辨率空间表示。该设计在不牺牲全局深度一致性的前提下保留了细结构与清晰边界。在各类零样本基准上,PXDepth兼具忠实的局部几何与有竞争力的全局深度精度,且推理效率高。我们的代码与模型可在该https URL获取。
英文摘要
Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-patch ViT encoders and convolutional decoders, as coarse tokenization can weaken pixel-level cues that upsampling cannot fully recover. To address this issue, we propose PXDepth, a discriminative monocular depth model that separates global context modeling from pixel-level depth prediction. Specifically, a large-patch ViT captures global scene context, while a pixel-space predictor composed of Context-Modulated Pixel Transformer blocks maintains high-resolution spatial representations throughout depth estimation. This design preserves fine structures and sharp boundaries without sacrificing global depth consistency. Across diverse zero-shot benchmarks, PXDepth combines faithful local geometry with competitive global depth accuracy while remaining efficient at inference. Our code and model are available at https://yuanzhy29.github.io/PXDepth-Page/.
发表机构
- CUHKSZ(香港中文大学(深圳))
- Shenzhen-FNii(深圳先进院集成所)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。