DDMS:将多视图基础特征判别式蒸馏到单视图模型中
DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models
浏览论文内容
中文总结 AI 辅助
本研究提出DDMS方法,通过判别式蒸馏将多视图模型的3D几何知识迁移到单视图模型,获得兼具3D一致性、局部判别性与语义迁移性的基础特征,在多维度实验中效果显著。
中文摘要 AI 辅助
DINO等基础视觉特征在现代计算机视觉中发挥着关键作用,近来已成为多视图前馈几何估计器的核心组件。本研究表明,通过将这些多视图模型(其内部的3D几何知识)重新蒸馏到单视图估计器中,可获得增强的3D一致性基础特征。核心思路是构建多视图教师模型,将预训练的2D基础特征与多视图几何特征融合,再用判别式排序目标优化融合表示。通过该判别式蒸馏框架,强制学习到的特征兼具3D一致性与局部判别性,同时保持与原始基础模型特征空间的对齐,以保留预训练表示的语义结构。一致性与局部判别性对跨图像形成语义和几何对应关系等3D计算机视觉问题至关重要。为验证方法有效性,开展了多维度综合实验:直接特征分析、密集预测迁移、显式3D提升与渲染。所有评估中,本方法持续生成更强的3D感知基础特征,在保留原始表示语义迁移性的同时,提升了多视图一致性与局部判别性。
英文摘要
Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models---their internal knowledge of 3D geometry---into a single-view estimator, we can obtain enhanced 3D consistent foundational features. Our key idea is to construct a multi-view teacher by fusing pretrained 2D foundation features with multi-view geometric features, and refining the fused representation with a discriminative ranking objective. Through our discriminative distillation framework, we enforce the learned features to be both 3D consistent and locally distinctive, while keeping them aligned with the feature space of the original foundation model to preserve the semantic structure of the pretrained representation. Consistency and local discriminability are critical for 3D computer vision problems such as forming semantic and geometric correspondences across images. To demonstrate the effectiveness of our method, we perform comprehensive experiments spanning multiple angles: direct feature analysis, dense prediction transfer, and explicit 3D lifting and rendering. Across these evaluations, our method consistently produces stronger 3D-aware foundation features that improve multi-view consistency and local discriminability while preserving the semantic transferability of the original representation.
发表机构
- University of British Columbia(不列颠哥伦比亚大学)
- Sony Semiconductor Solutions Corporation(索尼半导体解决方案公司)
- Sony Corporation(索尼公司)
机构由 AI 辅助整理,请以论文原文为准。