发表机构
Kyutai; Gradium(Kyutai; Gradium)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将强化学习应用于GLM-4-Voice语音模型,通过监督微调和可验证奖励的RL,无需额外推理标记即可提升口语数学推理准确性,结合流式推理达到74.8%的自由形式准确率,创下语音原生模型的新纪录。
AI 中文摘要
语音语言模型相较于级联系统,能够实现更丰富的人机口语交互,允许访问副语言信息并降低延迟。然而,它们在数学推理基准上的准确性落后于文本模型。带有可验证奖励的强化学习(RL)在扩展文本模型解决复杂问题和限制幻觉的能力方面发挥了重要作用。在这项工作中,我们探索将RL应用于GLM-4-Voice语音模型(Zeng等人,2024),以弥合文本和口语数学问题解决之间的差距。我们首先通过监督微调在合成的口语问答数据上使模型适应领域。然后我们表明,即使没有额外的推理标记,RL也能将GSM8K上的准确性提高到超过以往仅通过补充推理轨迹才能达到的语音模型水平。当与现有的流式推理技术结合时,我们进一步将自由形式准确性提高到74.8%。这为语音原生模型的口语数学能力建立了新的最先进水平。
英文摘要
Speech language models enable richer spoken interactions between humans and machines than cascaded systems, allowing access to paralinguistic information and lower latency. However, their accuracy on mathematical reasoning benchmarks has lagged behind those of text models. Reinforcement learning (RL) with verifiable rewards has been instrumental in extending text models' capabilities for solving complex problems and limiting hallucinations. In this work, we explore applying RL to the GLM-4-Voice speech model (Zeng et al., 2024) to bridge the gap between textual and spoken mathematical problem solving. We first adapt the model to the domain using supervised fine-tuning on synthesized spoken question-answering data. We then show that, even without extra reasoning tokens, RL improves the accuracy on GSM8K beyond levels previously achieved for speech models only with supplementary reasoning traces. When combined with existing streaming reasoning techniques, we show further gains to 74.8% free-form accuracy. This establishes a new state-of-the-art for mathematical spoken abilities with speech-native models.
CommentsAccepted at COLM 2026