arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RetroThinker:在语音大语言模型中实现回顾性思考

RetroThinker: Enabling Retrospective Thinking in Speech LLMs

Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed, David Harwath

arXiv 2609.11864首次发表:更新:

发表机构

The University of Texas at Austin; FAIR, Meta Superintelligence Labs(德克萨斯大学奥斯汀分校; Meta超级智能实验室基础人工智能研究部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RetroThinker通过多阶段后训练框架,使语音大语言模型在推理中动态修正思维链,在GSM8K上以可比延迟获得11%的绝对准确率提升。

AI 中文摘要

语音大语言模型(SpeechLLMs)提供了更低的延迟,并保留了通常在级联自动语音识别(ASR)和基于文本的LM架构中丢失的副语言细微差别。然而,在复杂推理任务上,它们仍然落后于纯文本LLM,而实时口语交互则施加了严格的延迟约束。尽管先前的工作采用思维链(CoT)和并发推理来增强推理能力而不引起过高的延迟,但固有的准确性-延迟权衡仍然存在。在本文中,我们研究流式SpeechLLM是否能够在推理过程中动态地修正其推理轨迹。我们引入了RetroThinker,一个多阶段的后训练框架,使Moshi模型能够在推理过程中进行自我验证和向前修正CoT步骤。RetroThinker结合了在精选的回顾性思考数据上的监督微调(SFT)和基于长度的直接偏好优化(DPO),以在早期推理期间(即用户在说话时并发推理)优化回顾性思考。在GSM8K基准上的评估表明,与非回顾性基线相比,RetroThinker显著改善了准确性-延迟权衡,在可比较的延迟下实现了11%的绝对准确性提升。

英文摘要

Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imposes strict latency constraints. Although prior works employ Chain-of-Thought (CoT) and concurrent reasoning to enhance reasoning capabilities without inducing prohibitive delays, an inherent accuracy-latency trade-off persists. In this paper, we investigate whether a streaming SpeechLLM can dynamically revise its reasoning traces on the fly. We introduce RetroThinker, a multi-stage post-training framework that equips the Moshi model to self-verify and forward-correct CoT steps during inference. RetroThinker combines supervised fine-tuning (SFT) on curated retrospective thinking data with length-based direct preference optimization (DPO) to optimize retrospective during early reasoning (i.e., reasoning concurrently while the user speaks). Evaluated on the GSM8K benchmark, RetroThinker significantly improves the accuracy-latency trade-off over non-retrospective baselines, achieving an 11% absolute accuracy gain at a comparable latency.

CommentsAccepted to IEEE SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑