发表机构
TTI-Chicago; Adobe; Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University; UC Berkeley(芝加哥丰田理工学院; 奥多比公司; 捷克技术大学捷克信息学、机器人学与控制论研究所; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视频蒙太奇挑战,提出FilmGPT自回归Transformer,通过在电影语料库训练捕捉电影“语法”,推理时用镜头约束解码算法选最佳镜头,在镜头预测和电影编辑任务中表现出色,还适用于多种视频蒙太奇应用。
AI 中文摘要
本文介绍了FilmGPT,一种自回归Transformer,旨在应对视频蒙太奇挑战,即将原始的、“不可观看”的镜头集合转变为连贯的电影序列。受现代语言模型语言学习启发,在大量电影语料库上训练长上下文自回归Transformer,直接从数据而非手工编码规则中隐式捕捉电影“语法”。推理时引入镜头约束解码算法,根据从电影中学到的统计模式从输入原始镜头中选择最佳下一个镜头。在镜头序列排序标准基准上用FilmGPT自回归模型进行下一个镜头预测,优于先前技术水平。通过用户研究在完整电影编辑任务上评估镜头约束解码算法,基于FilmGPT的编辑显著优于先前方法。最后展示了FilmGPT在视频蒙太奇广泛应用中的适用性。
英文摘要
This work introduces FilmGPT, an autoregressive transformer designed to address the challenge of video montage--turning a collection of raw, "unwatchable" footage into coherent cinematic sequences. Inspired by language learning in modern LLMs, we train a long-context autoregressive transformer on a large corpus of movies. The aim is to implicitly capture the "grammar" of film directly from data rather than from hand-coded rules. Unlike other generative models, FilmGPT does not generate any new video frames. Instead, at inference time, we introduce a footage-constrained decoding algorithm to select the best next shot from the input raw footage according to the statistical patterns learned from films. We first evaluate these learned statistics directly by using the FilmGPT autoregressive model for next shot prediction on a standard benchmark of shot sequence ordering, outperforming the previous state of the art. We then evaluate our footage-constrained decoding algorithm on the full film editing task via a user study, and find that our FilmGPT-based editing significantly outperforms previous approaches. Finally, we demonstrate the applicability of FilmGPT to a wide range of applications in video montage, from automatic video segment trimming to human-in-the-loop film editing.
Journal refSIGGRAPH Conference Papers '26, Article 129, 1-11, ACM, 2026