用于序列动作视频生成的关键帧锚定身份保留
Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation
浏览论文内容
中文总结 AI 辅助
针对身份保留文本到视频生成任务中主体身份保持和动作执行的挑战,提出无训练三阶段管道框架,包括动作感知提示优化、身份保留生成和身份感知推理增强,该方法在官方排行榜上排名第三,性能和通用性强。
中文摘要 AI 辅助
保留身份的文本到视频生成旨在合成一个能准确遵循文本描述且在整个过程中保持用户指定主体可识别性的视频。IPVG26挑战将此框架从单个整体提示扩展到时间结构化规范。模型还接收带时间戳的动作字幕序列,并按指定顺序呈现执行这些动作的主体。我们提出无训练的三阶段管道框架应对挑战。动作感知提示优化阶段将输入重写为指定每个动作终端状态的图像生成提示;身份保留生成阶段通过联合参考身份及其前一帧生成关键帧序列;身份感知推理增强阶段使用多参考引导和身份驱动噪声搜索合成中间段。我们的方法在官方Track 2排行榜上排名第三,展现了有竞争力的性能和强大的通用性。
英文摘要
Identity-preserving text-to-video generation aims to synthesize a video that accurately follows a textual description while maintaining the recognizability of a user-specified subject throughout. The IPVG26 challenge extends this framework from a single holistic prompt to a temporally structured specification. The model additionally receives a sequence of timestamped action captions and must render the subject performing these actions in the specified order. This temporal structure presents a challenge not encountered in previous identity-preserving generation tasks, as the subject must continuously perform a scripted sequence of distinct actions while maintaining a consistent identity. However, end-to-end video generators are prone to appearance drift as motion accumulates and the depicted actions change. We address this challenge with a training-free, three-stage pipeline framework. An action-aware prompt polishment stage first rewrites the inputs into image-generation prompts that specify the terminal state of each action. An identity-preserving generation stage then produces the keyframe sequence by conditioning each frame jointly on the reference identity and its predecessor, thereby decoupling time-invariant appearance from time-varying pose. Finally, an identity-aware inference enhancement stage synthesizes the intermediate segments using multi-reference guidance and identity-driven noise searching, both of which reinforce identity fidelity during sampling. Our method ranked third on the official Track 2 leaderboard, demonstrating competitive performance and strong generality.
发表机构
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。