arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27756cs.SDcs.CLcs.MMeess.AS

Cocktail-Talker:基于Turn Action GRPO的嘈杂社交环境下多说话人对话建模

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

Xilin Jiang, Riki Shimizu, Sukru Samet Dindar, Junkai Wu, Zhongweiyang Xu, Nima Mesgarani

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Cocktail-Talker框架,结合Turn Action GRPO,通过动作令牌建模助手行为,利用Cocktail-DialogGen生成数据,实现嘈杂社交环境下的多说话人口语对话建模,推动更自然的对话系统发展。

中文摘要 AI 辅助

口语对话系统通常针对干净的二元交互设计,仅包含单个用户与助手轮流发言。然而现实中的社交对话往往更复杂:多名说话人可能在无关语音和背景噪音中参与同一对话,每句话可能针对助手、其他说话人或完全无关。在此场景下,助手不仅要决定说什么,还要决定是否发言。本文提出Cocktail-Talker,这是一种用于嘈杂社交环境下多说话人口语对话建模的语音大语言模型框架,通过三个动作令牌对助手行为建模:<|respond|>(回应)、<|listen|>(倾听)和<|ignore|>(忽略),放置在回复或静音前。Cocktail-Talker通过监督微调与强化学习训练,以生成合适的动作令牌,且仅在<|respond|>模式下生成语音回复。为准备训练数据,开发了Cocktail-DialogGen,这是一种基于大语言模型的数据流水线,可模拟不同社交场景中具有说话人角色的真实多说话人对话。这些组件共同推动了能在复杂社交环境中更自然、有选择性交互的口语对话系统发展。

英文摘要

Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant speech and background noise. Each utterance may be directed to the assistant, addressed to another speaker, or completely irrelevant. In such settings, the assistant must decide not only what to say, but also whether to speak at all. In this paper, we introduce Cocktail-Talker, a speech LLM framework for multi-speaker spoken dialog modeling in noisy social environments. We model the assistant's behavior with three action tokens: <|respond|>, <|listen|>, and <|ignore|>, placed before a response or silence. Cocktail-Talker is trained via supervised finetuning and reinforcement learning to generate the appropriate action token and, only in <|respond|> mode, a speech response. To prepare the training data, we develop Cocktail-DialogGen, an LLM-based data pipeline that simulates realistic multi-speaker dialogs with speaker roles across diverse social settings. Together, these components take a step toward spoken dialog systems that interact more naturally and selectively in complex social environments.

↑