发表机构
Kling AI Research(可灵AI研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SubjectAnchor提出一种基于显式视觉记忆的主题感知记忆到视频生成范式,通过记忆构建、时间旋转位置编码和注意力分区,在多镜头叙事中保持主题身份与场景一致性,实验验证其跨镜头一致性优于基线。
AI 中文摘要
我们提出了SubjectAnchor,一种面向多镜头叙事的主题感知记忆到视频生成范式,其中当前镜头通过基于从先前镜头提取的显式视觉记忆进行条件生成。目标是保持镜头切换间的主题身份和场景一致性,同时保留逐镜头提示的可控性。基于Wan2.2-I2V-A14B构建,SubjectAnchor包含三个关键组件:主题相关记忆构建、主题感知的时间旋转位置编码以及记忆感知的注意力分区。对于每个目标镜头,该方法通过将每个所需主题追溯至其历史外观并检索最相关的预计算关键帧来构建紧凑的记忆库。这些记忆帧被编码为模型输入的显式视觉条件,同时不同主题被分配到分离的负时间槽以减少身份干扰。此外,记忆感知的注意力分区在共享骨干网络内调节记忆令牌与生成内容之间的交互。该公式保留了显式视觉记忆的外观锚定,同时与脚本驱动的逐镜头生成兼容。实验表明,SubjectAnchor在跨镜头身份一致性上优于代表性的基于记忆和整体式基线,同时保持有竞争力的视觉质量。
英文摘要
We present SubjectAnchor, a Subject-Aware Memory-to-Video paradigm for multi-shot storytelling in which the current shot is generated by conditioning on explicit visual memories extracted from previous shots. The objective is to preserve subject identity and scene consistency across cuts while retaining the controllability of shot-wise prompting. Built on Wan2.2-I2V-A14B, SubjectAnchor contains three key components: subject-related memory construction, subject-aware temporal rotary position encoding, and memory-aware attention partition. For each target shot, the method constructs a compact memory bank by tracing each required subject to its historical appearance and retrieving the most relevant precomputed keyframes. These memory frames are encoded into the model input as explicit visual conditions, while different subjects are assigned to separated negative temporal slots to reduce identity interference. In addition, memory-aware attention partition regulates the interaction between memory tokens and generated content within a shared backbone. This formulation preserves the appearance anchoring of explicit visual memory while remaining compatible with script-driven shot-by-shot generation. Experiments show that SubjectAnchor improves cross-shot identity consistency over representative memory-based and holistic baselines while maintaining competitive visual quality.
CommentsAccepted by ACM MM 2026