arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

揭示运动对电影摄影机轨迹的价值

Unveiling the Value of Motion for Cinematic Camera Trajectories

Ziqi Zhou, Yujian Yuan, Laura Sevilla-Lara

arXiv 2609.38683首次发表:更新:

发表机构

University of Edinburgh; The Hong Kong University of Science and Technology(爱丁堡大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现将摄影机轨迹表示为方向和速度优于传统姿态表示,提出CineGEN生成模型和CineScript数据集,提升轨迹-文本对齐与生成性能,并编码导演意图。

AI 中文摘要

电影摄影机运动是一种基本的叙事工具,它不仅由摄影机在场景中的位置决定,还由其运动的方向和速度决定。近期关于摄影机轨迹生成和与文本对齐的工作依赖于以姿态为中心的表示。虽然原则上网络可以推导出运动方向和速度,但我们发现在实践中这可能不会发生。事实上,在本文中,我们发现将摄影机轨迹表示从传统的逐帧姿态分解为方向和速度,在多个任务中带来了惊人的益处,包括轨迹到文本的对齐以及文本到轨迹的生成。为了准确评估前者,我们引入了一个简单且可靠的协议,克服了先前评估基线的局限性。对于后者,基于这一表示性见解,我们提出了一种新颖的摄影机轨迹生成模型CineGEN,该模型在多种指标上取得了优越的性能。我们还提出了一个新数据集CineScript,其中包含带有场景描述和更高级元数据的电影片段。这些新数据使我们能够测试模型捕获高级电影摄影信息的能力。我们表明,尽管表示简单,但通过方向和速度来表示摄影机轨迹不仅有助于在数值上实现更好的对齐和生成,而且本质上编码了复杂的导演意图。

英文摘要

Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation. To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGEN, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed not only helps numerically to achieve better alignment and generation, but also inherently encodes complex directorial intent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑