arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27123cs.CV

EditaLive! 面向直播的统一角色视频编辑

EditaLive! Unified Character Video Editing for Live Streaming

  • University of Macau(澳门大学)
  • vivo(vivo公司)
  • Great Bay University(大湾区大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun

AI总结:

EditaLive是适配直播的新型角色视频编辑框架,通过改进Wan-Animate模型,实现了低延迟、表情一致的实时角色视频编辑,达到了最优性能。

AI中文摘要:

传统视频编辑主要聚焦于场景级内容,而直播则更侧重人物主体。然而,将现有视频编辑方法直接应用于以人物为中心的直播仍具挑战性,因为这些方法可能会引入面部表情不一致问题,且通常依赖多个离线推理步骤,无法适配实时交互需求。我们提出EditaLive,这是一种用于实时直播角色视频编辑的新型框架。具体而言,我们从预训练的图像动画模型Wan-Animate出发,该模型天然将外观与运动解耦,我们通过参考帧编辑和基于收集的CharEdit-50K数据集进行视频重构,将其重新用作基于指令的以人物为中心的视频编辑的基础模型。此外,我们将模型从离线双向生成适配为因果流式生成,并设计了一种对齐的自回滚蒸馏策略,将模型压缩为两步采样器,其中固定RoPE和对齐强制减少训练-推理差异,首帧保留的稀疏注意力过滤冗余历史信息以缓解外观漂移。大量实验表明,EditaLive实现了最先进的编辑性能,同时忠实保留面部表情,并具备低延迟的实时流式推理能力。

英文摘要:

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.

↑