arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

记忆引导的从用户视频集生成B-roll序列

Memory-Guided B-Roll Generation from User Video Collections

Cusuh Ham, Fabian Caba Heilbron, Josef Sivic, Bryan Russell

arXiv 2610.01884首次发表:更新:

发表机构

Adobe Research; Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University(Adobe研究院; 捷克理工大学信息学、机器人与控制论研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MemComposer系统,利用用户视频集构建结构化记忆,规划、检索并生成有依据的B-roll序列,在提示遵循和视觉对齐上显著优于无依据方法。

AI 中文摘要

我们提出了一种基于集合的B-roll序列生成方法。给定用户的视频集、自然语言指令和目标时长,目标是生成一个多镜头序列,该序列补充用户的主要镜头(A-roll),同时保留集合中的角色、场景、物体和风格。此任务具有挑战性,因为必须从数小时的拍摄素材中选择视觉证据,以指导序列中每个镜头的生成。我们通过MemComposer应对这一挑战,这是一个三阶段系统,将原始素材转化为带有视觉参考(角色、场景、物体和风格)的结构化记忆,并利用它来规划、检索和生成有依据的B-roll序列。首先,在一次性离线阶段,MemComposer从原始视频构建以实体为中心的记忆。其次,它利用记忆和用户指令来规划有依据的序列,并为每个镜头检索条件帧。第三,它迭代地生成和批评序列,以强制身份、场景和序列级一致性。我们在用户偏好研究中评估了MemComposer,沿两个维度:提示遵循和与用户集合的视觉对齐。与无依据的文本到视频规划器相比,MemComposer在提示遵循上获胜60.0%,在视觉对齐上获胜92.8%,显示了集合记忆和参考检索的依据优势。与仅从拍摄素材组装的检索序列相比,MemComposer在提示遵循上获胜94.5%,显示了生成缺失镜头的价值,而仅检索序列在视觉对齐上被偏好58.2%的比较。

英文摘要

We introduce an approach for collection-grounded B-roll sequence generation. Given a user's video collection, a directive given in natural language, and a target duration, the goal is to produce a multi-shot sequence that complements the user's primary footage (A-roll) while preserving the collection's characters, settings, objects, and style. This task is challenging as one must choose the visual evidence from hours of captured footage that should guide the generation of each shot in the sequence. We address this challenge with MemComposer, a three-stage system that turns raw footage into a structured memory with visual references (characters, settings, objects, and style) and uses it to plan, retrieve, and generate grounded B-roll sequences. First, in a one-time offline stage, MemComposer constructs an entity-centric memory from raw video. Second, it uses the memory and user directive to plan a grounded sequence and retrieve conditioning frames for each shot. Third, it iteratively generates and critiques the sequence to enforce identity, setting, and sequence-level consistency. We evaluate MemComposer in a user preference study along two dimensions: prompt adherence and visual alignment to the user's collection. Against an ungrounded text-to-video planner, MemComposer wins 60.0\% of prompt-adherence and 92.8\% of visual-alignment comparisons, showing the grounding benefit of collection memory and reference retrieval. Against retrieval-only sequences assembled from captured footage, MemComposer wins 94.5\% of prompt-adherence comparisons, showing the value of generating missing shots, while retrieval-only sequences are preferred for visual alignment in 58.2\% of comparisons.

CommentsProject page at https://cusuh.github.io/MemComposer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑