arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于稀疏直接飞行时间传感器的密集度量深度补全

Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors

Hakyeong Kim, Ruicheng Wang, Chengtang Yao, Jiaolong Yang, Min H. Kim

arXiv 2608.04737首次发表:更新:

发表机构

KAIST; USTC; Microsoft Research Asia(韩国科学技术院; 中国科学技术大学; 微软亚洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种深度引导双分支视觉Transformer框架,结合dToF模拟流程,实现稀疏dToF到密集度量深度的零样本泛化补全,性能优于现有方法且高效。

AI 中文摘要

直接飞行时间(dToF)传感器可提供高精度的度量深度,在复杂真实场景中比间接ToF系统更具鲁棒性。然而,其高昂的制造成本和有限的光电二极管阵列尺寸会生成极度稀疏、低分辨率且含噪的深度图,无法满足VR/XR、机器人技术及3D感知等需要密集度量深度的任务需求。现有单目及深度补全方法难以处理dToF设备特有的采样模式和硬件伪影,在稀疏度高或噪声大时性能会显著下降。本文提出一种可泛化的框架,用于从稀疏dToF测量中补全密集度量深度,可适配不同传感器类型、稀疏度水平及噪声条件。模型采用深度引导的双分支视觉Transformer编码器,分别处理RGB图像和稀疏dToF测量值,同时通过掩码联合注意力模块使深度标记能可靠引导图像特征而不被覆盖。轻量解码器可高效重建密集度量深度,无需基于扩散或过度细化的后处理。为应对配对训练数据的稀缺问题,本文引入一套全面的dToF模拟流程,可复现闪光、亚VGA闪光及旋转传感器的特性,包括硬件导致的退化、不规则稀疏性及真实噪声分布。仅在合成数据上训练的模型,在6个数据集和3种真实dToF设备上实现了强零样本泛化,在精度和计算效率上均优于现有方法,为从稀疏直接ToF传感器补全密集度量深度提供了鲁棒且实用的解决方案,代码和模型已开源。

英文摘要

Direct Time-of-Flight (dToF) sensors provide highly accurate metric depth and are more robust than indirect ToF systems in challenging real-world conditions. However, their high manufacturing cost and limited photodiode array size produce depth maps that are extremely sparse, low-resolution, and noisy, making them unsuitable for VR/XR, robotics, and 3D perception tasks that require dense metric depth. Existing monocular and depth completion methods struggle to handle the unique sampling patterns and hardware artifacts of dToF devices, and their performance often deteriorates significantly under severe sparsity or noise. We present a generalizable framework for dense metric depth completion from sparse dToF measurements, capable of operating across diverse sensor types, sparsity levels, and noise conditions. Our model employs a depth-guided dual-branch Vision Transformer encoder that processes RGB images and sparse dToF measurements separately, while a masked joint attention module allows depth tokens to reliably guide image features without being overwritten by them. A lightweight decoder reconstructs dense metric depth efficiently, without diffusion-based or refinement-heavy post-processing. To address the scarcity of paired training data, we introduce a comprehensive dToF simulation pipeline that reproduces the characteristics of flash, sub-VGA flash, and rotating sensors, including hardware-induced degradation, irregular sparsity, and realistic noise distributions. Trained entirely on synthetic data, our model achieves strong zero-shot generalization across 6 datasets and 3 real dToF devices, outperforming state-of-the-art approaches in both accuracy and computational efficiency. This establishes a robust and practical solution for dense metric depth completion from sparse direct ToF sensors. Our code and models are open-sourced. See https://vclab.kaist.ac.kr/cvpr2026p3.

Journal refProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑