OmniMate:用于交互式虚拟化身的开放式实时流视听生成
OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
浏览论文内容
中文总结 AI 辅助
研究针对实时交互式流中生成范围未知和跨模态身份一致性退化的挑战,提出OmniMate框架,通过生成进度控制器和多参考条件模块,实现高质量、低延迟的开放式实时交互式视听虚拟化身生成及长期跨模态身份一致性。
中文摘要 AI 辅助
基于扩散的生成模型的进展为交互式虚拟化身系统提供了基础,但将这些模型扩展到实时交互式流仍具有挑战性。为此提出OmniMate框架,它能实时联合合成视觉内容、语音和音频效果,实现自然沉浸式多轮交互。引入生成进度控制器明确建模各流块生成进度,提出多参考条件模块保持长期跨模态身份一致性。实验表明OmniMate实现了高质量、低延迟流生成并保持长期视听一致性,支持多轮对话中的真实、连贯和响应式交互体验。
英文摘要
Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for interactive avatar systems. However, extending unified audio-visual synthesis to real-time interactive streaming remains challenging, as the generation horizon is unknown in advance and the generated identity may drift over long-term generation. To address these challenges, we propose OmniMate, a unified framework for open-ended real-time interactive audio-visual avatar generation. OmniMate jointly synthesizes visual content, speech, and sound effects in real time, enabling natural and immersive multi-turn interactions. To achieve adaptive response progression, we introduce a Generation Progress Controller (GPC) that explicitly models the generation progress of each streaming chunk, allowing the model to complete responses according to the desired progress and achieve seamless transitions between execution and listening states. To preserve long-term cross-modal identity consistency, we propose a Multi-Reference Conditioning Module (MRCM), which leverages multiple reference images and a reference speech segment to provide persistent visual and speaker identity cues throughout long-duration streaming interactions. Extensive experiments on an interaction-oriented adaptation of VerseBench demonstrate that OmniMate achieves high-quality, low-latency streaming generation while maintaining strong long-term audio-visual consistency. The results further show that OmniMate supports realistic, coherent, and responsive interactive avatar experiences over extended multi-turn conversations.
发表机构
- China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.(中国电信人工智能技术(北京)有限公司)
机构由 AI 辅助整理,请以论文原文为准。