arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Omni-LiveAvatar:分钟级实时流式视听联动数字人生成

Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar Generation

Lunjie Zhu, Xingtong Ge, Fangyu Lin, Yi Zhang, Zhening Liu, Mengfei Li, Yumeng Zhang, Guanglu Song, Yu Liu, Jun Zhang

arXiv 2608.13602首次发表:更新:

发表机构

Hong Kong University of Science and Technology; Vivix Group Limited(香港科技大学; Vivix集团有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Omni-LiveAvatar框架,通过渐进式自回归蒸馏等技术,实现分钟级实时流式视听联动数字人生成,速度较LTX-2提升33倍且性能优于基线模型。

AI 中文摘要

视听联动生成模型是沉浸式交互式数字人生成的基础,但现有多数模型依赖双向注意力与多步去噪,仅能生成短片段,无法支持长时间实时交互。本文提出首个分钟级实时流式视听联动数字人生成框架Omni-LiveAvatar,核心技术包括:1)渐进式自回归蒸馏流水线,将大型双向视听扩散模型LTX-2转换为无需辅助稳定机制的少步自回归生成器;2)同步视听长短时记忆模块,在有限内存预算下保持全局一致性;3)分层滚动提示规划策略,实现语义连贯演化与提示无缝过渡。大量实验表明,Omni-LiveAvatar可实时生成高质量、同步的分钟级数字人:在单块NVIDIA H200 GPU上,生成速度较其教师模型LTX-2提升33倍;在视觉质量、音频质量、跨模态同步性与人脸保真度等指标上,均优于加速基线模型。项目代码可访问指定URL获取。

英文摘要

Joint audio-video generative models serve as foundation for immersive and interactive digital-human generation. Nevertheless, most existing models rely on bidirectional attention and multi-step denoising and can generate only short clips, making them unsuitable for real-time interaction over extended durations. We present Omni-LiveAvatar, the first framework for minute-level, real-time streaming joint audio-video avatar generation. Specifically, we propose (1) a progressive autoregressive distillation pipeline that transfers a large bidirectional joint audio-video diffusion model into a few-step autoregressive generator without auxiliary stabilization mechanisms; (2) a synchronized audio-video long-short-term memory that preserves global consistency under a bounded memory budget; and (3) a hierarchical rolling prompt planning strategy that enables coherent semantic evolution and seamless prompt transitions. Extensive experiments show that Omni-LiveAvatar generates high-quality, synchronized minute-level avatars in real time. In terms of speed, it achieves a 33$\times$ generation speedup over its teacher, LTX-2, on a single NVIDIA H200 GPU; in terms of generation quality, it outperforms accelerated baselines across visual quality, audio quality, cross-modal synchronization, and human fidelity. Our code is available at https://github.com/Aoko955/Omni-LiveAvatar.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑