发表机构
Avaturn Live(Avaturn Live)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AVTR-1提出开源实时交互虚拟形象技术栈,含1.53亿参数流匹配运动生成器,实现同步音视频流,并引入R-DGG指标验证说话者语音对运动的预测贡献。
AI 中文摘要
说话头和双人交互模型现已实现实时推理,但仅靠快速运动生成并不能产生交互式对话。实时系统必须将模型输出与外部语音代理的语音同步,调度视频帧进行播放,并处理中断。我们提出AVTR-1,一个用于实时交互式虚拟形象对话的开源技术栈,其核心是一个紧凑的1.53亿参数自回归流匹配运动生成器,以双方参与者的音频为条件。我们通过自蒸馏将其音频编码器适配为流式处理。该技术栈将模型基于块状生成转化为由外部语音代理驱动的连续、同步音视频流,并解析推导了其对用户端延迟的贡献,且用两个商业语音代理验证了所得界限。进一步实验表明,AVTR-1在所有报告的视觉质量指标和大多数传统听讲运动指标上领先于对比的双人交互系统,同时在唇形同步方面保持竞争力。其推理运行时在数据中心和消费级GPU上均可实时运行。然而,传统听讲指标无法确定配对说话者的语音是否对生成的运动有贡献。因此,我们引入基于参考的定向格兰杰增益(R-DGG),用于衡量在考虑听者历史和说话者运动之后,说话者语音所携带的额外预测信息。R-DGG发现,对于录制的听者和所有评估的双人交互系统,存在统计上支持的预测依赖,但对于没有配对音频的说话头生成器或不匹配的说话者-听者对则不然。我们以组件特定许可证发布模型权重、渲染器和服务后端。
英文摘要
Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live system must synchronize the model's output with speech from an external voice agent, schedule video frames for playback, and handle interruptions. We introduce AVTR-1, an open stack for real-time interactive avatar conversations, built around a compact 153M-parameter autoregressive flow-matching motion generator conditioned on both participants' audio. We adapt its audio encoder for streaming through self-distillation. The stack turns the model's chunk-based generation into a continuous, synchronized audio-video stream driven by an external voice agent, and we analytically derive its contribution to the user-facing latencies and validate the resulting bounds with two commercial voice agents. Further experiments demonstrate that AVTR-1 leads the compared dyadic systems on all reported visual-quality metrics and most conventional listening-motion metrics while remaining competitive in lip synchronization. Its inference runtime operates in real time on data-center and consumer GPUs. However, conventional listening metrics do not establish whether the paired speaker's speech contributes to generated motion. We therefore introduce the Reference-Based Directed Granger Gain (R-DGG), which measures the additional predictive information carried by speaker speech after accounting for listener history and speaker motion. R-DGG finds statistically supported predictive dependence for recorded listeners and all evaluated dyadic systems, but not for talking-head generators without paired audio or mismatched speaker-listener pairs. We release the model weights, renderer, and serving backend under component-specific licenses.