AI 中文总结
PIVOT是用于研究真实世界3D重建中姿态、内参等因素的多轨迹数据集与测试平台,其基准测试揭示了新视点合成方法在不同轨迹、姿态及内参下的性能差异。
AI 中文摘要
神经辐射场(NeRFs)、3D高斯溅射(3DGS)及相关新视点合成方法,通常在比机器人、无人机和自主系统实际遇到的更干净的捕获与重建条件下进行评估。基准测试常依赖于利于重建的轨迹、经优化的相机姿态与内参,以及从训练期间使用的轨迹中采样的保留视图。这些假设可能掩盖使用实测姿态、可复用相机标定以及结构不同的相机路径时的性能。我们推出PIVOT(Pose, Intrinsics and Viewpoint Oriented Testbed,即面向姿态、内参与视点的测试平台),这是一个用于独立研究这些因素的多轨迹数据集、处理流水线和评估框架。PIVOT使用不同的相机轨迹捕获每个场景,并在可用时保留两种姿态:传感器派生的实测姿态和COLMAP优化的姿态,以及标定和优化后的相机内参。它定义了三类基准系列:(1)所见与未见轨迹的新视点泛化;(2)实测与优化姿态的敏感性;(3)标定与优化内参的敏感性。我们还引入了定向姿态空间的Chamfer距离,以量化训练姿态对评估轨迹的覆盖程度。PIVOT v1包含5个用DJI Mini 4 Pro捕获的真实世界场景,并提供开放的处理流程和基于Nerfstudio的评估工具链。基准测试结果显示,在代表轨迹的保留视图与未见视图之间存在一致的质量差距,且对姿态源和相机内参存在显著敏感性。
英文摘要
Neural radiance fields (NeRFs), 3D Gaussian Splatting (3DGS), and related novel-view synthesis methods are commonly evaluated under capture and reconstruction conditions cleaner than those encountered by robots, drones, and autonomous systems. Benchmarks often rely on reconstruction-friendly trajectories, optimized camera poses and intrinsics, and held-out views sampled from trajectories represented during training. These assumptions can obscure performance with measured poses, reusable camera calibration, and structurally different camera paths. We introduce PIVOT (Pose, Intrinsics and Viewpoint Oriented Testbed), a multi-trajectory dataset, processing pipeline, and evaluation framework for independently studying these factors. PIVOT captures each scene using diverse camera trajectories and retains, where available, both sensor-derived measured poses and COLMAP-optimized poses, together with calibrated and optimized camera intrinsics. It defines three benchmark families: (1) seen versus unseen trajectory novel-view generalization, (2) measured versus optimized pose sensitivity, and (3) calibrated versus optimized intrinsics sensitivity. We also introduce a directed pose-space Chamfer distance to quantify how well training poses cover an evaluation trajectory. PIVOT v1 contains five real-world scenes captured with a DJI Mini 4 Pro and provides an open processing and Nerfstudio-based evaluation toolchain. Benchmark results show a consistent quality gap between held-out views on represented trajectories and unseen trajectories, as well as substantial sensitivity to pose source and camera intrinsics.