发表机构
Fudan University; Shanghai Innovation Institute; Tencent Hunyuan; Zhejiang University(复旦大学; 上海创新研究院; 腾讯混元; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OmniCam提出一种基于几何锚定位姿标记学习的自回归模型,从全景图和文本生成相机轨迹,在OmniCaT基准上显著降低轨迹误差和碰撞率。
AI 中文摘要
相机轨迹控制着视频生成、场景重建和机器人感知中的视点变化。从语言生成轨迹需要同时考虑场景几何和目标感知取景。我们提出了OmniCam,一种自回归模型,能够从单张全景图和文本轨迹描述中生成相机位姿序列。其几何锚定位姿标记学习结合了三个组成部分:用于全向几何上下文的全景点云编码器;具有时间一致四元数符号的混合绝对旋转和相对平移标记化;以及带有显式3D目标锚点的分离的几何和语义条件流。我们还构建了OmniCaT,包含267,700条轨迹,涵盖四种相机行为。在报告的OmniCaT评估中,与在OmniCaT上重新训练的GenDoP相比,OmniCam将轨迹误差降低了28%至47%,碰撞率降低了65.8%。针对每个指标的最佳基线,ATE和碰撞率分别降低了43.0%和62.3%。组件消融实验支持了几何和目标感知条件的使用,而下游实验则考察了相机控制的视频生成和机器人主动感知。
英文摘要
Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28--47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.