AI 中文总结
JoyAI-Talker采用Thinker-Talker架构等技术,实现兼具强推理能力与共情交互的全双工语音对话系统,在基准测试及用户打断、背景语音场景下表现优异。
AI 中文摘要
我们提出了JoyAI-Talker,这是一个全双工语音对话系统,既具备强大的基础模型能力,又能实现共情交互与语音智能体智能。JoyAI-Talker采用模块化的Thinker-Talker架构,并进一步实施统一的语音-文本联合训练流程,以缓解常见的“认知退化”瓶颈,从而在将模型核心的文本推理、STEM(科学、技术、工程、数学)及逻辑能力扩展至基于语音的交互的同时,大幅保留这些能力。针对表达性语音合成,Talker模块采用文本可控的生成范式,支持通过自然语言指令灵活控制语音属性与局部副语言事件(如笑声、叹息),以实现更具表达性和细粒度的语音响应。为提升对话共情能力,我们引入了Persona-Adaptive Empathetic Response(PAER,角色自适应共情响应)框架。PAER采用分层认知流程,从原始输入音频中提取性别、年龄、情绪状态等非言语说话者线索,将其融入Thinker的CoT(思维链)推理中,生成语义适配的响应,该响应需在语义合适的文本与话语层面表达性、局部副语言事件(包括叹息、语速、音量)的细粒度控制之间实现平衡。我们还集成了Joy-Duplex,这是一种状态驱动、即插即用的全双工框架,可作为高效门控引擎用于实时轮次控制。大量评估表明,JoyAI-Talker在基础T2T(文本到文本)和S2T(语音到文本)基准上取得了极具竞争力的性能;在全双工评估中,该系统在用户打断下达到0.88的高响应率,同时在背景语音下保持极低的误触发率,证明其已具备流畅自然的语音对话能力。
英文摘要
We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.