arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AptAvatar:用于生产就绪头像的快速且生动的长格式音频驱动视频生成

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

Hengyuan Zhang, Jingna Sun, Meiguang Jin, Junfeng Ma

arXiv 2607.24013首次发表:更新:

AI 中文总结

研究生产就绪的音频驱动头像生成,提出AptAvatar框架,通过端点锚定分布蒸馏和自生成历史重放方法,实现快速且生动的长格式视频生成,仅用2次归一化流评估,加速60倍且保持视觉保真度和长跨度身份。

AI 中文摘要

生产就绪的音频驱动头像生成需要在不牺牲保真度或运动表现力的情况下进行高效推理。然而,现有加速方法往往通过因果注意力和短时间跨度等限制性架构选择,或通过降低模型容量和分辨率来牺牲质量。为此,我们提出了AptAvatar,一个拥有14B参数的长格式音频驱动头像生成框架,可实现快速且富有表现力的推理。为提高生产级应用的效率,AptAvatar应对极端的两步生成挑战。我们引入端点锚定分布蒸馏来弥合多步教师模型和两步学生模型之间的差距,还引入自生成历史重放来提高长跨度一致性。大量实验表明,AptAvatar仅用2次归一化流评估就能生成生动的720p长格式头像视频,在保持视觉保真度和长跨度身份的同时实现了60倍的加速。

英文摘要

Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often compromise quality through restrictive architectural choices, such as causal attention and short temporal horizons, or by reducing model capacity and resolution. Without such compromises, we propose AptAvatar, a 14B-parameter long-form audio-driven avatar generation framework that delivers fast and expressive inference. For efficiency in production-level applications, AptAvatar addresses the extreme two-step generation challenge. To bridge the gap between the multi-step teacher model and the two-step student model, we introduce Endpoint-Anchored Distribution Distillation. It augments vanilla distribution matching with a dedicated Anchor Score Estimator trained on the trajectory-endpoint distribution defined from a frozen pretrained 4-step bridge generator. This provides an attainable endpoint-level anchor for the evolving two-step student. To improve long-horizon consistency, we further introduce Self-Generated History Replay, which reuses cached outputs from earlier generator checkpoints as history conditions during chunk-wise training. This approximates inference-time conditioning on self-generated histories without costly online rollouts, mitigating quality degradation from accumulated history errors. Extensive experiments demonstrate that AptAvatar generates vivid 720p long-form avatar videos with only 2 NFEs, achieving a 60x speedup while preserving visual fidelity and long-horizon identity. Code is available at https://github.com/TaoLiveAIGC/AptAvatar

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑