发表机构
University of Rochester; Adobe Research(罗切斯特大学; Adobe 研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ALIVE通过首帧编辑和指令,使插入物体与源视频内容产生连贯交互,并训练VLM预测引导,显著提升交互保真度。
AI 中文摘要
当前的视频编辑器可以插入物体,但往往难以使这些物体参与交互,例如被拿起或操控。我们提出了ALIVE,一个通过连贯的交互使插入的物体“活起来”的框架,该框架利用编辑后的首帧和仅指明新增物体的指令,使插入物体与源视频内容进行交互。我们整理了35,800个编辑对,结合了3D渲染、模型生成和真实世界视频,以及来自ROSE的通用编辑对。每个编辑对在目标物体是否存在上有所不同,同时保留周围动作,从而教会编辑器协调的物体行为和源内容保持。我们进一步训练了一个视觉语言模型(VLM),从相同输入预测交互引导。我们引入了ALIVE交互基准,使用统一的基于VLM的协议评估交互保真度、源内容保持和视觉连贯性,并在通用视频物体插入基准上进行评估。在没有VLM引导的情况下,ALIVE在两个基准上分别比最强基线高出43.9%和4.4%。VLM预测的引导进一步将ALIVE交互分数提高了0.95分,且无需额外用户输入。
英文摘要
Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects "alive" through coherent interactions with the source video's contents, using an edited first frame and an instruction naming only the added object. We curate 35,800 editing pairs combining 3D-rendered, model-generated, and real-world videos with general editing pairs from ROSE. Each pair differs in the target object's presence while preserving the surrounding action, teaching editors coordinated object behavior and source preservation. We further train a vision-language model (VLM) to predict interaction guidance from the same inputs. We introduce the ALIVE-interaction benchmark to assess interaction fidelity, source preservation, and visual coherence using a unified VLM-based protocol, and evaluate on the general video object insertion benchmark. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% and 4.4% on the two benchmarks, respectively. VLM-predicted guidance further improves the ALIVE-interaction score by 0.95 points without additional user inputs.
CommentsProject page: https://real-time-video-research.github.io/alive/