发表机构
The Chinese University of Hong Kong, Shenzhen; Tencent Hunyuan; Tsinghua University; The Hong Kong University of Science and Technology; Amphion Technology Co., Ltd.(香港中文大学(深圳); 腾讯混元; 清华大学; 香港科技大学; Amphion科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AURAL通过潜在空间分布建模与联合分块预测,在保持CoT级推理性能的同时将首词延迟降低11.8倍,并构建AuralReason-683K数据集与自适应推理预算机制。
AI 中文摘要
模型智能与快速响应共同决定语音语言模型的交互质量,但二者难以兼得。显式思维链(CoT)能提升推理与音频理解能力,但生成中间推理标记会延迟响应。描述细粒度声学线索会进一步加长思维链并增加延迟。潜在推理可减少这一开销,但现有方法往往落后于CoT,且受限于单路径监督和不适配问题难度的推理预算。我们提出AURAL,在潜在空间中对多条合理推理延续的分布进行建模,并联合预测未来状态的分块,以减少顺序前向传播和推理延迟。为给潜在推理提供初始监督,我们构建了AuralReason-683K:包含68.3万条双语语音话语(约1000小时),附带用于情感识别、共情对话和通用推理的精简CoT。AURAL-RL随后在超越这些轨迹的基础上进行探索,奖励产生高质量答案的精简推理,并根据每个问题调整推理努力。在两个骨干网络上,AURAL-RL取得了与CoT-RL相当的性能,在大多数指标上相较于各自的监督检查点有更大提升。分析进一步表明,更难的问题会引发更多的潜在推理步骤。在Qwen2.5-Omni上,它将首个答案标记的时间从1.22秒降至0.10秒,缩短了11.8倍,而直接回答为0.05秒。
英文摘要
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.