Omni 交互代理技术报告
Multimodal Duplex Interaction Agent
浏览论文内容
中文总结 AI 辅助
Gander 是一个端到端模型,通过小脑-大脑协作和流式思考者-说话者架构,统一全模态感知、实时交互与代理能力,实现全双工交互,并展现出竞争力的性能与鲁棒性。
中文摘要 AI 辅助
在这项工作中,我们提出了 Gander,一个端到端模型,在单一框架内统一了全模态感知、实时交互和代理能力。与基于回合的传统范式不同,Gander 持续接收跨多种模态的流式输入,包括视频、语音和文本,从而在日常对话和复杂工作流导向的代理场景中实现自然的全双工交互。用户可以随时打断模型,而模型也可以主动提供中间反馈或提出后续问题。为了原生支持这些能力,Gander 采用了两个关键架构设计:1)它采用小脑-大脑协作框架,其中小脑负责实时交互和全模态对话能力,而大脑处理复杂推理和更高级的代理任务。这两个组件通过工具调用和代理编排运行时持续交互。2)小脑构建在流式思考者-说话者架构之上,用户输入和模型输出在块级别进一步展平为有序的令牌流,为低延迟、持续交互提供统一表示。我们从四个维度对 Gander 进行了全面评估:对话能力、全模态理解、交互能力和代理智能。内部人工评估表明,Gander 保持了最先进开源模型的自然且富有表现力的口语对话能力,同时在全模态交互方面取得了具有竞争力的性能。Gander 还在具有挑战性的现实场景中表现出鲁棒性,包括背景噪声干扰、多方交互和反馈通道通信。我们发布了 Gander 及其模型、代码和数据,以促进社区的进一步研究和发展。
英文摘要
In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtime interaction with an asynchronous agent loop. In contrast to conventional turn based systems, Gander continuously processes streaming user inputs, enabling full-duplex interaction in both everyday conversations and complex workflow agent scenarios. Users can interrupt an ongoing response, while the model can proactively provide intermediate feedback or ask follow up questions. To natively support these capabilities, Gander adopts two key architectural designs: 1) a Cerebellum-Brain collaborative framework, Cerebellum is responsible for realtime interaction while the Brain handles complex reasoning and higher level agentic tasks. The two components interact continuously through tool calling and the agent orchestration runtime. 2) The Cerebellum is built upon a streaming Thinker-Talker architecture, where user inputs and model outputs are flattened into an ordered token stream at the chunk level. We evaluate Gander across conversational ability, interactive capability, understanding, and tool assisted task execution. Internal human evaluations show that Gander maintains natural and expressive spoken dialogue, while benchmark results demonstrate effective turn taking capability and encouraging results on spoken question answering and related understanding tasks. Gander also supports a range of challenging interaction settings, including background noise interference, multi-party interactions, and backchannel communication. While our current evaluation focuses on tool assisted settings, broader long horizon agent tasks and more diverse deployment conditions remain promising directions for further study. We release Gander together with its models, code, and data to facilitate further research and development in the community.
发表机构
- Tencent(腾讯)
- Zhejiang University(浙江大学)
- Shanghai Jiao Tong University(上海交通大学)
- The Chinese University of Hong Kong(香港中文大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。