arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KeyID:用于身份保留视频生成的解耦草稿生成与关键帧编辑

KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation

Jianjie Luo, Yiming Zhong, Haoming Shen, Yupeng Xiao, Zhenguo Yang

arXiv 2608.16154首次发表:更新:

发表机构

Guangdong University of Technology(广东工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KeyID是无需训练的IPVG框架,通过解耦视频动态合成与身份注入,解决身份一致性问题,在ACM MM 2026 IPVG挑战赛获亚军,可扩展至多主体参考及复杂序列动作生成。

AI 中文摘要

身份保留视频生成(IPVG)需要合成既忠实于参考主体又符合文本提示的视频。现有方法常受限于高调优成本或有限的输入级增强,在复杂长序列动作中难以维持严格的身份一致性。为解决这些局限,我们提出KeyID,这是一种无需训练的IPVG框架,它将视频动态合成与身份注入解耦。具体而言,KeyID包含两个组件:(1)参考感知视频生成,生成与多个参考对齐的身份无关视频草稿;(2)身份保留关键帧编辑,通过稀疏关键帧校正及后续运动插值整合目标身份。通过从密集帧级监督转向稀疏关键帧级细化,KeyID有效解决了提示词贴合度与身份保真度之间的容量冲突。关键的是,其模块化设计可无缝扩展至多主体参考及复杂序列动作生成,无需额外训练。KeyID在官方挑战基准的自动与人工评估中优于现有工作,最终获得ACM多媒体2026 IPVG挑战赛第2赛道(序列动作)的亚军。源代码可在指定网址获取。

英文摘要

Identity-preserving video generation (IPVG) requires synthesizing videos that are faithful to both reference subjects and text prompts. Existing methods are often hindered by high tuning costs or limited input-level enhancements, struggling to maintain rigid identity consistency during complex, long-sequence actions. To address these limitations, we propose KeyID, a training-free IPVG framework that decouples the synthesis of video dynamics from the injection of identity. Specifically, KeyID comprises two components: (1) Reference-Aware Video Generation, which produces an identity-agnostic video draft aligned with multiple references, and (2) Identity-Preserved Keyframe Editing, which integrates the target identity via sparse keyframe correction and subsequent motion interpolation. By shifting from dense frame-level supervision to sparse keyframe-level refinement, KeyID effectively resolves the capacity conflict between prompt adherence and identity fidelity. Crucially, our modular design allows seamless extension to multi-subject references and complex sequential action generation without additional training. KeyID outperforms prior works and is validated by automatic and human evaluations on the official challenge benchmark, ultimately securing the runner-up position in the Track 2 (Sequential Action) of the ACM Multimedia 2026 IPVG Grand Challenge. Source code is available at https://github.com/WISLab-GDUT/KeyID.

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑