arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JoyAI-Talker:面向共情语音智能体的全双工语音交互大模型

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

Yinhao Bai, Jinming Chen, Yafeng Chen, Wei Deng, Boya Dong, Nan Duan, Yu Gu, Weisheng Han, Yankun Huang, Ming Ke, Hao Li, Jingdong Li, Xiangyu Liang, Ning Liu, Yuan Liu, Ji Miao, Jiaqi Wang, Qi Wang, Wenchao Wang, Yuxuan Wang, Zhenfang Wang, Zhangyu Xiao, Chao Xue, Hongfei Xue, Fan Yu, Tianyi Zhang, Yuan Zhang, Yuqi Zhang, Lin Zhu

arXiv 2608.01119首次发表:更新:

AI 中文总结

JoyAI-Talker采用Thinker-Talker架构等技术,实现兼具强推理能力与共情交互的全双工语音对话系统,在基准测试及用户打断、背景语音场景下表现优异。

AI 中文摘要

我们提出了JoyAI-Talker,这是一个全双工语音对话系统,既具备强大的基础模型能力,又能实现共情交互与语音智能体智能。JoyAI-Talker采用模块化的Thinker-Talker架构,并进一步实施统一的语音-文本联合训练流程,以缓解常见的“认知退化”瓶颈,从而在将模型核心的文本推理、STEM(科学、技术、工程、数学)及逻辑能力扩展至基于语音的交互的同时,大幅保留这些能力。针对表达性语音合成,Talker模块采用文本可控的生成范式,支持通过自然语言指令灵活控制语音属性与局部副语言事件(如笑声、叹息),以实现更具表达性和细粒度的语音响应。为提升对话共情能力,我们引入了Persona-Adaptive Empathetic Response(PAER,角色自适应共情响应)框架。PAER采用分层认知流程,从原始输入音频中提取性别、年龄、情绪状态等非言语说话者线索,将其融入Thinker的CoT(思维链)推理中,生成语义适配的响应,该响应需在语义合适的文本与话语层面表达性、局部副语言事件(包括叹息、语速、音量)的细粒度控制之间实现平衡。我们还集成了Joy-Duplex,这是一种状态驱动、即插即用的全双工框架,可作为高效门控引擎用于实时轮次控制。大量评估表明,JoyAI-Talker在基础T2T(文本到文本)和S2T(语音到文本)基准上取得了极具竞争力的性能;在全双工评估中,该系统在用户打断下达到0.88的高响应率,同时在背景语音下保持极低的误触发率,证明其已具备流畅自然的语音对话能力。

英文摘要

We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑