发表机构
Snap Inc.; King Abdullah University of Science and Technology (KAUST)(Snap公司; 阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究第一人称视角视频流编辑问题,提出EgoPlay,通过微调预训练模型,在单个端到端模型中联合学习事件识别等,构建数据集训练双向视频扩散编辑器,在Ego4D基准测试中表现出色,优于现有基线。
AI 中文摘要
我们介绍了EgoPlay,一种用于第一人称视角视频流的事件触发式视频到视频编辑器,它通过在主要从Ego4D构建的事件条件数据上微调预训练的V2V扩散变换器获得。给定一个单目视频和一个“当X发生时,做Y”形式的事件触发提示,EgoPlay推断事件X是否以及何时发生,保留事件前的帧,并仅对事件后的延续应用编辑Y。EgoPlay不是将单独的事件检测器与编辑器级联,而是在单个端到端模型中联合学习事件识别、时间约束和像素级编辑,同时还处理否定和多事件提示。为了支持这一点,我们构建了一个包含106K个事件触发剪辑-提示对的大规模数据集,涵盖正触发、伪造触发否定和多事件提示。然后,我们训练了一个具有事件触发监督的双向视频扩散编辑器,并导出了一个用于逐块可流式推理的因果变体。我们还引入了一种事件感知评估协议,分别测量触发后编辑质量、触发前保留和误触发鲁棒性。在Ego4D基准测试中,EgoPlay显著优于EgoEdit,这是基于指令的第一人称视角视频编辑的当前最先进基线,在编辑质量、视觉质量和背景一致性方面的相对增益分别为17.7%、16.9%和16.4%。在相同指标上,它也比VLM引导的检测器-编辑器基线高出15.7%、14.5%和13.5%,同时使用的GPU内存不到一半。
英文摘要
We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form "when X happens, do Y," EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the post-event continuation. Rather than cascading a separate event detector with an editor, EgoPlay learns event recognition, temporal restraint, and pixel-level editing jointly in a single end-to-end model, while also handling negative and multi-event prompts. To support this, we construct a large-scale dataset of 106K event-triggered clip-prompt pairs spanning positive triggers, fabricated-trigger negatives, and multi-event prompts. We then train a bidirectional video diffusion editor with event-triggered supervision and derive a causal variant for chunk-by-chunk streamable inference. We further introduce an event-aware evaluation protocol that separately measures post-trigger editing quality, pre-trigger preservation, and false-trigger robustness. On the Ego4D benchmark, EgoPlay substantially outperforms EgoEdit, the state-of-the-art instruction-based egocentric video editing baseline, with relative gains of 17.7%, 16.9%, and 16.4% in editing quality, visual quality, and background consistency. It also surpasses a VLM-guided detector-editor baseline by 15.7%, 14.5%, and 13.5% on the same metrics, while using less than half the GPU memory.
CommentsAccepted to SIGGRAPH Asia 2026 as a Conference Paper. Project page: https://egoplay2026.github.io/egoplay