arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TKCAM:基于文本与关键帧的相机轨迹生成

TKCAM: Text and Keyframe to Camera Trajectory Generation

Haozhe Yang, Zhiyang Dou, Zekai Gu, Cheng Lin, Wenping Wang, Yuan Liu, Taku Komura

arXiv 2610.11105首次发表:更新:

发表机构

The University of Hong Kong; The Hong Kong University of Science and Technology; Macau University of Science and Technology; Texas A&M University(香港大学; 香港科技大学; 澳门科技大学; 德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TKCAM是基于生成式掩码建模的文本与关键帧条件相机运动合成框架,通过RVQ和两阶段掩码Transformer实现轨迹生成,在FID等指标上优于SOTA,还构建了相关数据集与基准。

AI 中文摘要

生成高质量且可控的相机运动对AI辅助电影制作、视频合成及3D场景理解至关重要。我们提出TKCAM,这是一种基于生成式掩码建模的文本与关键帧条件相机运动合成框架。我们用包含位置、速度和连续旋转表示的12维运动学特征来表示相机动力学,并通过残差向量量化器(RVQ)将其离散化为分层运动令牌。随后,两阶段掩码Transformer架构学习重建和细化这些令牌,利用显式自注意力和交叉注意力模块实现多模态条件控制。该框架的核心特征是稀疏视觉关键帧条件控制:用户可提供自由形式的文本提示,以及选定时间戳的RGB观测值,这些观测值为生成连贯的中间轨迹提供时间局部化的视觉指导。此外,为推进评估标准,我们整理了大规模文本-相机数据集RealEstate10K-Cap,并通过通用CLaTr评估器建立了跨域基准。大量实验表明,TKCAM在Fréchet距离(FID)、文本-运动匹配分数及检索指标(R@K)上均优于近期的SOTA基线,额外分析还评估了时间平滑性和跨域泛化能力。代码可在该https URL获取。

英文摘要

Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Fréchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization. Code is available at https://github.com/linearalgebrayhz/TKCAM.

Comments20 pages, 7 figures. Paper accepted to NeurIPS2026 (submission number 15461)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑