发表机构
Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有商业视频生成模型的身份漂移等缺陷,提出AESR轻量级框架,含智能体提示增强与视觉语义修复模块,搭配混合专家选择策略,其系统MIPL_Video在ACM MM 2026挑战赛赛道1获第一
AI 中文摘要
身份保留视频生成旨在合成既遵循自然语言指令、又能保持给定主体视觉身份的视频。近期的商业视频生成模型已实现出色的视觉质量与运动真实性,但仍存在身份漂移、指令遵循不完整、复杂提示下视觉细节缺失等问题。由于这些模型通常是闭源黑箱,通过参数优化直接改进它们往往不可行。因此,我们提出Agentic Enhancement and Semantic Repair(AESR,智能体增强与语义修复),这是一个用于身份保留视频生成的轻量级增强框架。为改进生成前的提示构建并缓解上述缺陷,AESR引入了全局智能体提示增强模块:该模块从官方文档学习特定模型的提示格式,从人机交互数据获取以人类为中心的视频生成先验,并通过智能体循环将测试域的身份保留生成经验积累为可复用的手册。为进一步修复经增强提示生成的视频中的错误,AESR引入了样本级视觉语义修复模块:该模块使用视觉语言模型(VLM)定位视频中的错误片段并设计修复指令,将选定帧编辑为明确的视觉参考,再引导视频编辑模型修复局部语义或身份相关错误。我们还采用轻量级混合专家(Mixture-of-Experts)选择策略,从不同生成与优化路径中选择可靠输出。在ACM MM 2026身份保留视频生成挑战赛的官方评估协议下,我们的系统MIPL_Video在赛道1中排名第一,证明了AESR在实际身份保留视频生成中的有效性。代码可在该https链接获取。
英文摘要
Identity-preserving video generation aims to synthesize videos that follow natural-language instructions while maintaining the visual identity of a given subject. Recent commercial video generation models have achieved strong visual quality and motion realism, but they still suffer from identity drift, incomplete instruction following, and missing visual details under complex prompts. Since these models are usually closed-source black boxes, directly improving them through parameter optimization is often infeasible. We therefore propose Agentic Enhancement and Semantic Repair (AESR), a lightweight enhancement framework for identity-preserving video generation. To improve prompt construction before generation and mitigate the above failures, AESR introduces a global agentic prompt enhancement module. This module learns model-specific prompting formats from official documentation, acquires human-centered video generation priors from human-interaction data, and accumulates test-domain identity-preserving generation experience into a reusable playbook through an agentic loop. To further repair errors in videos generated with enhanced prompts, AESR introduces a sample-level visual semantic repair module, which uses a VLM to locate erroneous video segments and design repair instructions, edits selected frames into explicit visual references, and guides a video editing model to fix local semantic or identity-related errors. We also adopt a lightweight Mixture-of-Experts selection strategy to choose reliable outputs from different generation and refinement paths. Under the official evaluation protocol of the ACM MM 2026 Identity-Preserving Video Generation Challenge, our system MIPL\_Video ranked first in Track 1, demonstrating the effectiveness of AESR for practical identity-preserving video generation. The code is available at https://github.com/oceanflowlab/AESR.