LiveAnimate:稳定的长时长流驱动人体动画实时生成系统
LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time
浏览论文内容
中文总结 AI 辅助
LiveAnimate基于14B参数视频DiT,经两阶段训练结合PR-Sink等设计,实现19.63 FPS长时长流人体动画生成,兼顾实时性、质量与稳定性,优于现有系统。
中文摘要 AI 辅助
由姿态驱动的人体动画合成任务是从单张参考图像和驱动姿态流中生成目标人物的视频,实时生成是直播、远程呈现、虚拟化身等交互式应用的核心需求,但基于扩散模型的系统每段剪辑需耗时数分钟至数小时,无法实现响应式交互。本文提出LiveAnimate,据所知是首个结合实时流生成与十亿级参数长时长稳定生成的动画系统,基于140亿参数的视频扩散Transformer(DiT)构建。两阶段训练流程:首先通过参考锚定教师强制适配将预训练双向DiT改造为块因果自回归生成器,再通过块式自强制蒸馏将采样预算降至3步。为在长流中保留外观信息,本文提出姿态检索汇注意力(PR-Sink),这是一种受限键值缓存机制,包含永久锚定首个生成块的静态汇、存储姿态检索历史块的动态汇,以及三槽滚动窗口;当姿态重复时,PR-Sink可恢复相关外观上下文,无需保留整个序列,因此内存和每块延迟不随流时长变化。结合Ulysses序列并行与算子融合,这些设计在2块NVIDIA H100 GPU上实现19.63 FPS的流推理。在3分钟基准测试中,LiveAnimate从最初30秒到最后1分钟保持几乎恒定的感知质量和身份(IQA分别为4.047和4.026),而现有系统要么大幅降质,要么需耗时数小时离线计算才能完成相同时长生成。这些结果为交互式全身动画建立了质量、延迟和时长的新平衡点。
英文摘要
Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose stream. Real-time generation is essential for interactive applications such as live streaming, telepresence, and virtual avatars, yet diffusion-based systems require minutes to hours per clip, precluding responsive interaction. We present LiveAnimate, to our knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT). A two-stage training pipeline first adapts a pretrained bidirectional DiT into a block-causal autoregressive generator through Reference-Anchored Teacher-Forcing Adaptation, and then reduces the sampling budget to three steps through Block-wise Self-Forcing Distillation. To preserve appearance over extended streams, we introduce Pose-Retrieval Sink Attention (PR-Sink), a bounded KV-cache mechanism combining a Static Sink that permanently anchors the first generated block, a Dynamic Sink that holds a pose-retrieved historical block, and a three-slot Rolling Window. When a pose recurs, PR-Sink restores the relevant appearance context without retaining the entire sequence, so memory and per-block latency remain constant regardless of stream duration. Together with Ulysses sequence parallelism and operator fusion, these designs enable 19.63\,FPS streaming inference on two NVIDIA H100 GPUs. On a three-minute benchmark, LiveAnimate maintains nearly constant perceptual quality and identity from the first 30 seconds to the final minute, while prior systems degrade substantially or require hours of offline computation for the same rollout. These results establish a new operating point in quality, latency, and duration for interactive full-body animation.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- Qwen Applications Business Group of Alibaba(阿里巴巴通义千问应用业务组)
- Liblib AI
机构由 AI 辅助整理,请以论文原文为准。