arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EgoPlay:用于第一人称视角视频流的事件触发式视频编辑

EgoPlay: Event-Triggered Video Editing for Egocentric Streams

Jinjie Mai, Gordon Guocheng Qian, Willi Menapace, Arpit Sahni, Chaoyang Wang, Ashkan Mirzaei, Runjia Li, Sergey Tulyakov, Bernard Ghanem, Peter Wonka, Rameen Abdal

arXiv 2607.24560首次发表:更新:

发表机构

Snap Inc.; King Abdullah University of Science and Technology (KAUST)(Snap公司; 阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究第一人称视角视频流编辑问题,提出EgoPlay,通过微调预训练模型,在单个端到端模型中联合学习事件识别等,构建数据集训练双向视频扩散编辑器,在Ego4D基准测试中表现出色,优于现有基线。

AI 中文摘要

我们介绍了EgoPlay,一种用于第一人称视角视频流的事件触发式视频到视频编辑器,它通过在主要从Ego4D构建的事件条件数据上微调预训练的V2V扩散变换器获得。给定一个单目视频和一个“当X发生时,做Y”形式的事件触发提示,EgoPlay推断事件X是否以及何时发生,保留事件前的帧,并仅对事件后的延续应用编辑Y。EgoPlay不是将单独的事件检测器与编辑器级联,而是在单个端到端模型中联合学习事件识别、时间约束和像素级编辑,同时还处理否定和多事件提示。为了支持这一点,我们构建了一个包含106K个事件触发剪辑-提示对的大规模数据集,涵盖正触发、伪造触发否定和多事件提示。然后,我们训练了一个具有事件触发监督的双向视频扩散编辑器,并导出了一个用于逐块可流式推理的因果变体。我们还引入了一种事件感知评估协议,分别测量触发后编辑质量、触发前保留和误触发鲁棒性。在Ego4D基准测试中,EgoPlay显著优于EgoEdit,这是基于指令的第一人称视角视频编辑的当前最先进基线,在编辑质量、视觉质量和背景一致性方面的相对增益分别为17.7%、16.9%和16.4%。在相同指标上,它也比VLM引导的检测器-编辑器基线高出15.7%、14.5%和13.5%,同时使用的GPU内存不到一半。

英文摘要

We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form "when X happens, do Y," EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the post-event continuation. Rather than cascading a separate event detector with an editor, EgoPlay learns event recognition, temporal restraint, and pixel-level editing jointly in a single end-to-end model, while also handling negative and multi-event prompts. To support this, we construct a large-scale dataset of 106K event-triggered clip-prompt pairs spanning positive triggers, fabricated-trigger negatives, and multi-event prompts. We then train a bidirectional video diffusion editor with event-triggered supervision and derive a causal variant for chunk-by-chunk streamable inference. We further introduce an event-aware evaluation protocol that separately measures post-trigger editing quality, pre-trigger preservation, and false-trigger robustness. On the Ego4D benchmark, EgoPlay substantially outperforms EgoEdit, the state-of-the-art instruction-based egocentric video editing baseline, with relative gains of 17.7%, 16.9%, and 16.4% in editing quality, visual quality, and background consistency. It also surpasses a VLM-guided detector-editor baseline by 15.7%, 14.5%, and 13.5% on the same metrics, while using less than half the GPU memory.

CommentsAccepted to SIGGRAPH Asia 2026 as a Conference Paper. Project page: https://egoplay2026.github.io/egoplay

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑