arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OTT3R:以1%计算量实现多视图3D重建与快速数据集生成

OTT3R: Multi-View 3D Reconstruction and Fast Dataset Generation at 1% Compute

Brandon Leblanc, Charalambos Poullis

arXiv 2609.36374首次发表:更新:

发表机构

Concordia University(康考迪亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OTT3R通过知识蒸馏将大型3D重建模型压缩至1%计算量,实现快速推理,并提供高吞吐伪标签生成替代COLMAP,在7-Scenes上准确率提升4倍。

AI 中文摘要

前馈式3D重建模型通过扩大模型和数据集规模取得了令人瞩目的性能,但其成本将大多数研究团队拒之门外,并阻碍了边缘部署。此外,在没有传感器的情况下生成3D监督仍然依赖于缓慢且不可靠的运动恢复结构(Structure-from-Motion),因为社区缺乏一个类似COLMAP的神经3D伪标签生成系统。我们提出了OTT3R(仅RGB的小型Transformer用于3D重建),这是一个知识蒸馏框架,可在配备2块GPU的单台工作站上解决这两个问题。将π³(9.59亿参数)蒸馏为1.02亿参数的学生模型,实现了9.4倍压缩和高达7倍的推理加速,训练计算量仅为VGGT的1.6%。集成的伪标签流水线为COLMAP提供了一种可靠、高吞吐量的替代方案,在3.5小时内于两块商用GPU上为包含66.7万张图像的语料库生成了密集的逐像素点图和SE(3)相机位姿,并在我们测试的每个序列上都取得成功,包括COLMAP失败的序列。通用学生模型在分布内的单目深度估计上紧随教师模型,并在零样本情况下在7-Scenes和DTU补全上优于COLMAP,但在分布外的多视图几何上无法取代教师模型。可部署的成果是领域特化的学生模型:在0.2%计算量下完成特化后,其在7-Scenes上的准确率比COLMAP高4倍,吞吐量高980倍,且补全质量接近教师模型。代码可在以下网址获取:https://this https URL

英文摘要

Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks a COLMAP-like system for neural 3D pseudo-label generation. We present OTT3R (RGB-Only Tiny Transformer for 3D Reconstruction), a knowledge distillation framework that addresses both problems on a single workstation equipped with 2 GPUs. Distilling $π^3$ (959M parameters) into a 102M-parameter student yields 9.4$\times$ compression and up to 7$\times$ faster inference, trained at 1.6% of VGGT's training compute. An integrated pseudo-label pipeline offers a reliable, high-throughput alternative to COLMAP, generating dense per-pixel point maps and SE(3) camera poses for a 667K-image corpus in 3.5 hours on two commodity GPUs and succeeding on every sequence we tested, including those where COLMAP fails. The general student tracks the teacher on in-distribution monocular depth and, zero-shot, outperforms COLMAP on 7-Scenes and on DTU completion, but it does not replace the teacher on out-of-distribution multi-view geometry. The deployable artifact is the domain-specialized student: after specialization at 0.2% compute, it is 4$\times$ more accurate than COLMAP on 7-Scenes at 980$\times$ throughput, with near-teacher completion. Code is available at https://github.com/TheFourthKaramazov/OTT3R

CommentsAccepted at ACCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑