LiteMVS:结合基础模型蒸馏与专家聚合的高效多视图立体匹配
LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation
浏览论文内容
中文总结 AI 辅助
LiteMVS是集成平面扫描几何推理与单目先验的轻量多视图深度模型,通过MoE和基础模型蒸馏提升重建质量与效率,在ScanNetv2等数据集上表现优异。
中文摘要 AI 辅助
实时三维感知对机器人、增强现实和具身智能应用至关重要。现有多视图立体(MVS)方法主要依赖几何对应关系,在无纹理或重复区域常失效;而单目深度模型虽利用了强大的图像级先验,却缺乏稳健的多视图几何约束。更重要的是,在机器人和具身操作场景中,高质量三维几何不仅对静态重建必不可少,还是学习时间一致的四维表示的关键基础。为获得更强结构感知、更具时空扩展潜力的视觉表示,我们提出LiteMVS,一种集成平面扫描几何推理与强大单目语义及结构先验的轻量多视图深度估计模型。LiteMVS的核心思路是将轻量分割模型和大规模视觉基础模型得到的高级单目知识高效注入多视图立体框架。具体而言,LiteMVS用语义描述符丰富代价体积,并采用混合专家(MoE)公式实现深度假设间的自适应几何聚合;此外,从视觉基础模型蒸馏得到的几何先验进一步增强单目引导,且不增加推理成本。通过该设计,LiteMVS不仅提升了静态场景的深度估计和三维重建质量,还为后续时间建模和四维表示学习提供了更可靠的几何基础。在ScanNetv2和7-Scenes上的实验表明,LiteMVS在保持竞争力效率的同时,实现了高质量的深度预测和三维重建。
英文摘要
Real-time 3D perception is crucial for robotics, augmented reality, and embodied intelligence applications. Existing multi-view stereo (MVS) methods primarily rely on geometric correspondences, which often fail in textureless or repetitive regions, while monocular depth models leverage strong image-level priors but lack robust multi-view geometric constraints. More importantly, in robotics and embodied manipulation scenarios, high-quality 3D geometry is not only essential for static reconstruction, but also serves as a critical foundation for learning temporally consistent 4D representations. To obtain visual representations with stronger structural awareness and greater potential for spatiotemporal extension, we present LiteMVS, a lightweight multi-view depth estimation model that integrates plane-sweep geometric reasoning with strong monocular semantic and structural priors. The central idea of LiteMVS is to efficiently inject high-level monocular knowledge, obtained from lightweight segmentation models and large-scale vision foundation models, into a multi-view stereo framework. In particular, LiteMVS enriches the cost volume with semantic descriptors and employs a Mixture-of-Experts (MoE) formulation to enable adaptive geometric aggregation across depth hypotheses. Moreover, geometric priors distilled from vision foundation models further strengthen monocular guidance without increasing inference cost. Through this design, LiteMVS not only improves depth estimation and 3D reconstruction quality in static scenes, but also provides a more reliable geometric foundation for subsequent temporal modeling and 4D representation learning. Experiments on ScanNetv2 and 7-Scenes demonstrate that LiteMVS achieves high-quality depth prediction and 3D reconstruction while maintaining competitive efficiency.