发表机构
The Hong Kong Polytechnic University; ByteDance; AMD(香港理工大学; 字节跳动; 超威半导体公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Avatar-Forever是一种解耦并行训练框架,通过双分支设计与ForeverCache机制,在单张H100 GPU上实现27.2 FPS的768x512音频驱动无限虚拟形象生成,保障质量与效率,推进稳定数字人研究。
AI 中文摘要
现有流视频系统通常依赖以蒸馏为核心的顺序训练流水线来实现少步长长视频生成,但该范式存在两大局限:其一,早期阶段出现的故障或分布偏移会影响后续优化,使训练收敛过程复杂化;其二,以蒸馏为核心的目标函数偏向短期生成,但在长序列推演中自回归误差累积时易出现质量下降。我们提出Avatar-Forever,这是一个用于高质量实时无限交互虚拟形象的解耦并行训练框架。我们不再在顺序蒸馏流水线中耦合生成效率与长视界鲁棒性,而是将其视为两个可并行训练的独立能力:一个分支执行全参数蒸馏以训练具备高视觉质量的高效生成器,另一个分支通过面向恢复的推演训练(Recovery-oriented Rollout Training,RRT)训练轻量型长视界适配器,以提升长视界推理条件下的生成鲁棒性。我们的解耦并行训练设计简化了整体训练过程,避免了少步长生成与长视界适配之间不必要的目标冲突。我们进一步引入ForeverCache,一种分块特征缓存机制,以大幅减少流推理过程中的冗余历史计算。Avatar-Forever基于22B视频基础模型构建,支持无边界的音频驱动虚拟形象生成,同时保持身份一致性、运动连贯性与视觉保真度,在单张H100 GPU上可实现高分辨率768x512视频的端到端吞吐量达27.2 FPS,为稳定数字人提供了可行路径。
英文摘要
Existing streaming video systems often rely on sequential, distillation-centered training pipelines to enable few-step long-video generation. However, this paradigm suffers from two limitations. First, failures or distribution shifts introduced in earlier stages affect later optimization, complicating the training process to converge. Second, the distillation-centric objective favours short-term generation but is prone to quality degradation when autoregressive errors accumulate over long rollouts. We propose Avatar-Forever, a decoupled parallel training framework for high-quality real-time infinite interactive avatars. Instead of coupling generation efficiency and long-horizon robustness under a sequential distillation pipeline, we treat them as two independent capabilities that can be trained in parallel. One branch performs full-parameter distillation to train an efficient generator with high visual quality, while another trains a lightweight long-horizon adapter via Recovery-oriented Rollout Training (RRT), which improves generation robustness under long-horizon inference conditions. Our decoupled parallel training design simplifies the overall training process and avoids unnecessary objective conflicts between few-step generation and long-horizon adaptation. We further introduce ForeverCache, a chunk-wise feature caching mechanism to substantially reduce redundant history computation during streaming inference. Built upon a 22B video foundation model, Avatar-Forever supports unbounded audio-driven avatar generation while maintaining identity consistency, motion coherence, and visual fidelity, enabling an end-to-end throughput of high-resolution 768x512 videos at 27.2 FPS on a single H100 GPU and providing a practical path toward stable digital humans.