发表机构
Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对动态相机内参估计的训练数据匮乏、评测基准多样性不足问题,构建含大规模合成数据集与真实基准的InFlux++,可提升RGB图像的逐帧内参估计性能。
AI 中文摘要
相机内参对于从2D视频中恢复3D结构至关重要。然而,大多数3D算法假设视频全程内参固定,这一假设在真实野外视频中往往不成立。因此,从RGB图像估计逐帧内参对于让3D方法对含动态内参的视频具备鲁棒性至关重要。此前InFlux通过首个带逐帧内参真值的动态内参真实视频基准推动了该方向研究,但现有方法仍受两个问题制约:(i)训练数据稀缺且内参多样性不足;(ii)包括InFlux在内的基准场景与相机运动多样性有限,难以充分评估方法性能。为解决这两个缺口,本文提出InFlux++,它包含两个组件。InFlux++ Synth是大规模过程式生成的合成视频数据集,含1841个高分辨率视频的44.1万余标注帧,为动态内参预测模型训练提供精准的逐帧内参真值,部分子集还包含逐帧位姿、深度与法向量。该类视频通过相机变焦与对焦变化实现丰富的内参多样性,同时包含动态对象及镜头畸变、散焦模糊等真实渲染效果。InFlux++ Real是大规模真实世界基准,在InFlux基础上新增334个高分辨率视频的51.4万余采集帧,覆盖更广泛的场景与相机运动。在InFlux++ Synth上对现有内参预测方法进行微调,可在InFlux++ Real与InFlux数据集上持续提升焦距估计精度,表明合成监督在基于RGB的内参预测任务中具备应用前景。数据集、基准、代码、视频、提交指南及实时排行榜可访问对应https URL获取。
英文摘要
Camera intrinsics are vital for recovering 3D structure from 2D video. However, most 3D algorithms assume fixed intrinsics throughout a video, an assumption that often fails for real-world in-the-wild videos. Consequently, estimating per-frame intrinsics from RGB images is critical for making 3D methods robust to videos with dynamic intrinsics. InFlux previously advanced this research direction by establishing the first real-world benchmark with per-frame ground truth intrinsics for dynamic intrinsics videos. Nevertheless, existing methods remain inaccurate due to two obstacles: (i) training data is scarce and lacks intrinsics diversity; and (ii) benchmarks, including InFlux, have limited scene and camera motion diversity, making it difficult to properly evaluate methods. To address both gaps, we present InFlux++, consisting of two components. InFlux++ Synth is a large-scale procedurally generated synthetic video dataset with 441K+ annotated frames from 1841 high-resolution videos, providing accurate per-frame ground truth intrinsics for training dynamic intrinsics prediction models; a subset also includes per-frame pose, depth, and normals. The videos feature rich intrinsics diversity through changes in camera zoom and focus, as well as dynamic objects and realistic rendering effects such as lens distortion and defocus blur. InFlux++ Real is a large-scale real-world benchmark that extends InFlux with 514K+ newly captured frames across 334 high-resolution videos, spanning a wider range of scenes and camera motions. Finetuning existing intrinsics prediction methods on InFlux++ Synth consistently improves focal length estimation across both InFlux++ Real and InFlux, suggesting that synthetic supervision is promising for RGB-based intrinsics prediction. For the dataset, benchmark, code, videos, submission instructions, and live leaderboard, please visit https://influx.cs.princeton.edu/.
CommentsAccepted to ECCV 2026