arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LVSPM:长序列视图合成与位姿估计模型

LVSPM: Long Sequence View Synthesis and Pose Estimation Model

Xi Chen, Yachi Zhang, Linghao Chen, Minghua Liu, Hao Su, Zexiang Xu, Xiaoshuai Zhang

arXiv 2610.10960首次发表:更新:

发表机构

UC San Diego; Sudo AI GmbH(加州大学圣地亚哥分校; Sudo AI 有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LVSPM是一种可泛化模型,联合估计相机位姿并合成新视图,采用测试时训练层,在多数据集的位姿估计和新视图合成任务中超越VGGT等基线,性能随场景规模增长仍稳定。

AI 中文摘要

我们提出了LVSPM,这是一种可泛化的模型,能从未校准的图像集合中联合估计相机位姿并合成新视图。仅使用RGB图像和位姿监督进行训练,LVSPM无需密集的3D真值,且采用测试时训练(TTT)层,可无缝扩展至数百个输入视图。在RealEstate10k、Co3Dv2和DL3DV数据集上,LVSPM在16至256个视图的位姿估计任务中超越了VGGT,在严格阈值下优势尤为显著。在更视图覆盖更大场景的实用协议下的新视图合成任务中,LVSPM实现了无位姿依赖的最优质量——其峰值信噪比(PSNR)甚至超越了依赖位姿的模型,且在场景规模增长时仍能保持高质量,而基线模型则出现性能崩溃。代码可在该https URL获取。

英文摘要

We present LVSPM, a generalizable model that jointly estimates camera poses and synthesizes novel views from uncalibrated image collections. Trained with only RGB images and pose supervision, LVSPM avoids dense 3D ground truth and employs test-time training (TTT) layers to scale seamlessly to hundreds of input views. On RealEstate10k, Co3Dv2, and DL3DV, LVSPM surpasses VGGT in pose estimation across 16-256 views, with especially large margins at strict thresholds. For novel view synthesis under a practical protocol where more views cover larger scenes, LVSPM achieves state-of-the-art pose-free quality---surpassing even pose-dependent models in PSNR---and still maintains high quality as scene scale grows, while baselines collapse. The code is available at https://burningdust21.github.io/Projects/LVSPM .

CommentsECCV 2026. Project Page: https://burningdust21.github.io/Projects/LVSPM/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑