arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度思考,直接表达:面向副语言接地口语对话的循环潜在推理

Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

Shengbo Cai, Yuxiang Wang, Jingran Xie, Zhisheng Zhang, Shun Lei, Di Cao, Teddy Sun, Zhiyong Wu

arXiv 2609.37818首次发表:更新:

发表机构

Tencent Hunyuan(腾讯混元)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对共情口语对话中感知与推理脱节的问题,提出LoopSLM,利用循环Transformer进行潜在推理,以声学接地细化隐藏状态,无需思维链,在提升推理准确率的同时大幅降低延迟和令牌数。

AI 中文摘要

共情口语对话要求模型同时利用说话的内容和说话的方式来决定如何回应。显式思维链(CoT)可以改善副语言感知,并使回复中的声学线索更加明确,但不能确保这些线索在回复规划中被有效使用。我们将这种不匹配称为感知-推理差距。此外,思维链可能无法完全捕捉词语中的声学线索,并且生成思维链会增加推理延迟。为解决这些局限,我们提出了LoopSLM,它基于循环Transformer进行潜在推理,在每次迭代中重用解码器块,通过声学接地来细化隐藏状态。其两阶段训练通过将学习推理与学习回应分离,进一步缩小了感知-推理差距,使得无需思维链即可直接推理。在EchoMind上,LoopSLM在副语言理解、推理和回复质量方面优于Qwen2.5-Omni-7B。与CoT-SFT基线相比,LoopSLM在推理准确率上提升了超过20个百分点,同时生成的令牌数减少了64.5%,延迟减半。它还在大多数共情回复指标上优于Qwen3-Omni-Thinking,延迟降低了34倍。尽管仅使用对话数据进行训练,LoopSLM在通用音频基准上的准确率也有所提升。

英文摘要

Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acoustic grounding at every pass. Its two-stage training further narrows the perception-reasoning gap by separating learning to reason from learning to respond, enabling direct inference without CoT. On EchoMind, LoopSLM improves paralinguistic understanding, reasoning, and reply quality over Qwen2.5-Omni-7B. Against the CoT-SFT baseline, LoopSLM gains over 20 points in reasoning accuracy while generating 64.5% fewer tokens at half the latency. It also outperforms Qwen3-Omni-Thinking on most empathetic reply metrics with 34x lower latency. Despite training only on dialogue data, LoopSLM improves accuracy on general audio benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑