能够出声思考的语音语言模型
Spoken Language Models that Think Aloud
查看机构详情
- Meta Superintelligence Labs(Meta超级智能实验室)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对语音语言模型在串行推理中产生静默的问题,提出异步出声思考框架,通过双流协调减少静默并保持准确性。
中文摘要 AI 辅助
虽然思维链(CoT)推理提升了语言模型的能力,但将其直接应用于语音语言模型(SLMs)时,在串行的“先思考后说话”范式下可能会引入较长的静默间隔,从而干扰实时语音交互。为解决这一问题,我们提出了一种在“思考者-说话者”架构内用于基于推理的语音语言模型的异步出声思考框架。该框架维护一个用于逻辑推理的主推理流,以及一个轻量级的出声思考流,该流根据用户输入和不断演化的推理状态生成简短的、基于任务进展的语句。一种动态平衡策略在运行时协调这两个流,触发额外的出声思考语音以避免静默间隙,并在最终回答准备就绪时取消待处理的语句。在语音推理和问答基准上的实验表明,我们的方法在推理过程中大幅减少了用户可感知的静默,同时保持了与串行“先思考后说话”基线相当的答案准确性,展示了异步出声思考在语音语言模型中实现响应式交互的潜力。
英文摘要
While Chain-of-Thought (CoT) reasoning has improved the capability of language models, directly applying it to Spoken Language Models (SLMs) may introduce long silent intervals under the serial "think-then-speak" paradigm, disrupting real-time spoken interaction. To address this issue, we propose an asynchronous think-aloud framework for reasoning-based SLMs within the Thinker-Talker architecture. The framework maintains a primary reasoning stream for logical deduction and a lightweight think-aloud stream that generates short, task-grounded progress utterances conditioned on the user input and the evolving reasoning state. A dynamic balance strategy coordinates the two streams at runtime, triggering additional think-aloud speech to avoid silent gaps and canceling pending utterances when the final response becomes ready. Experiments on spoken reasoning and question-answering benchmarks show that our approach substantially reduces user-audible silence during reasoning while maintaining answer accuracy comparable to that of a serial "think-then-speak" baseline, demonstrating the potential of asynchronous think-aloud for responsive interaction in SLMs.