SFE-VGGT:基于事件的单目深度估计的无源VGGT蒸馏
SFE-VGGT: Source-Free VGGT Distillation for Event-Based Monocular Depth Estimation
浏览论文内容
中文总结 AI 辅助
提出SFE-VGGT无源蒸馏框架,无需配对RGB数据,通过替代帧教师和可靠性感知蒸馏将VGGT几何先验迁移至事件域,夜间场景精度显著提升。
中文摘要 AI 辅助
近期基于事件的深度估计方法通过跨模态蒸馏成功地将视觉基础模型的几何先验迁移到事件域。然而,这些方法在训练过程中依赖同步的RGB-事件对或深度标注,严重限制了实际部署。为克服这一瓶颈,我们提出SFE-VGGT,一种新颖的无源框架,在没有任何配对RGB观测的情况下,将VGGT的几何先验蒸馏到事件域。我们的核心思想是直接从目标事件流重建替代帧,作为冻结的几何教师,完全消除对真实源RGB数据的需求。关键的是,由于这些替代帧固有地产生不完美且空间变化的监督,直接从中蒸馏会传播伪影。为解决此问题,我们引入一种新颖的可靠性感知蒸馏策略,包括密度感知特征蒸馏以强调信息丰富的事件区域,以及置信度加权深度蒸馏以基于教师-学生相对预测置信度动态调节监督。同时,我们提出跨帧关系一致性损失,利用可靠的帧间对应强制时间几何稳定性,无需时间一致的教师深度。大量实验表明,尽管无源,我们的SFE-VGGT在标准条件下与依赖RGB的基线精度紧密匹配,并在具有挑战性的夜间场景中显著超越它们。在MVSEC夜间序列中,与EventVGGT相比,SFE-VGGT将平均10米深度误差降低15.3%。此外,我们的方法在真实世界数据集上展现出鲁棒的零样本泛化能力,证明高度有效的几何先验可以在严格无源监督下迁移到事件相机。
英文摘要
Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However, their reliance on synchronized RGB-event pairs or depth annotations during training severely restricts practical deployment. To overcome this bottleneck, we propose SFE-VGGT, a novel source-free framework that distills the geometric priors of VGGT to the event domain without any paired RGB observations. Our core idea is to reconstruct surrogate frames directly from the target event stream to act as a frozen geometric teacher, entirely eliminating the need for genuine source RGB data. Crucially, as these surrogate frames inherently yield imperfect and spatially varying supervision, directly distilling from them propagates artifacts. To resolve this, we introduce a novel reliability-aware distillation strategy. This includes Density-Aware Feature Distillation to emphasize informative event regions, and Confidence-Weighted Depth Distillation to dynamically regulate supervision based on relative teacher-student prediction confidence. Meanwhile, we propose a Cross-Frame Relational Consistency loss that enforces temporal geometric stability using reliable inter-frame correspondences, bypassing the need for temporally consistent teacher's depth. Extensive experiments demonstrate that, despite source-free, our SFE-VGGT closely matches the accuracy of RGB-dependent baselines under standard conditions and significantly surpasses them in challenging nighttime scenarios. Across MVSEC nighttime sequences, SFE-VGGT reduces the average 10 m depth error by 15.3% compared with EventVGGT. Moreover, our method exhibits robust zero-shot generalization across real-world datasets, proving that highly effective geometric priors can be transferred to event cameras using strictly source-free supervision.
发表机构
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。