Realtime-Venus:一种具有异步委托的全双工交互系统
Realtime-Venus: A full-duplex interaction system with asynchronous delegation
浏览论文内容
中文总结 AI 辅助
提出Realtime-Venus全双工交互系统,采用两个9B模型和异步委托机制,在视频和音频基准上取得领先性能,支持实时交互与后台任务并行。
中文摘要 AI 辅助
数字与物理环境中的自然交互需要持续的感知和及时的响应。口语对话依赖于声学和语言线索,而视频交互还需要将对话锚定在不断演变的视觉上下文中。我们提出了Realtime-Venus,一个主动式全双工交互系统,包含两个分别训练的9B模型:用于音视频交互的Realtime-Venus-Omni和用于口语交互的Realtime-Venus-Audio。每个模型都作为一个完整的对话前端,通过共享的用户输入、模型输出和委托事件因果时间线,整合持续感知、对话控制和原生语音生成。一个双循环运行时协调实时交互与后台推理和工具执行。前台交互继续进行,同时Realtime-Venus-Harness异步执行任务并返回结果以整合到正在进行的对话中。两个模型都遵循一个共同的训练后流程,结合了离线理解、主动全双工轨迹和委托工作流。在评估的在线模型中,Realtime-Venus-Omni在八个视频基准中的六个上取得了最高分,包括StreamingBench(70.2%)、OVO-Bench(64.7%)和Daily-Omni(81.3%)。在八个音频理解和口语问答基准中,Realtime-Venus-Audio在MMAU(78.0%)、MMAU-Pro(63.2%)、Llama Questions(83.8%)和Speech CMMLU(67.8%)上领先于对比模型,同时匹配了最佳的VoiceBench AlpacaEval得分4.81。在Full-Duplex-Bench v1.5上,Realtime-Venus-Audio对75%的用户打断做出响应,并在反馈语、面向他人的语音和背景语音下分别实现了97%、88%和86%的继续率,在全部三个继续指标上均超过了Gemini 3.1 Live和GPT-4o。
英文摘要
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events. A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue. Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows. Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.
发表机构
- Ant Group(蚂蚁集团)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。