发表机构
Leibniz Research Center, Huawei; The Chinese University of Hong Kong; City University of Hong Kong(华为莱布尼茨研究中心; 香港中文大学; 香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语音语言模型中S2TS模式与S2T模式间的输出模式差距,提出联合输出在策略蒸馏方法,利用教师软目标显著降低准确率差距,实验验证其有效性。
AI 中文摘要
自回归生成交错文本和声学标记是语音大语言模型中生成口语响应的常见方法。尽管这种设计能够实现带有显式文本指导的流式生成,但生成的声学标记会成为后续文本预测上下文的一部分。在给定相同语音输入的情况下,我们观察到在语音到文本和语音(S2TS)模式下生成的内部文本的答案准确率明显低于语音到文本(S2T)响应。我们将这种差异称为“输出模式差距”(OMG)。为缩小OMG,我们提出了“联合输出在策略蒸馏”(JO-OPD),该方法利用学生生成的S2TS轨迹,将模型更强的S2T策略蒸馏到联合生成中。在每个文本位置,S2T教师从学生先前输出的纯文本投影提供软目标,而学生则从相应的完整交错历史中进行预测。一个保留目标进一步正则化非文本预测。在Step-Audio-2-mini和Baichuan-Audio-Instruct上的实验揭示了两种交错生成架构中的OMG。在Step-Audio-2-mini上,JO-OPD将Spoken-MQA上的OMG从42.87个百分点降至16.26个百分点,将语音渲染的GSM8K上的OMG从29.72个百分点降至13.04个百分点,同时S2T准确率变化很小,且降幅显著大于匹配的SFT基线。基于ASR的评估进一步显示,在Spoken-MQA上口语答案准确率提高了7.49个百分点。
英文摘要
Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models. Although this design enables streaming generation with explicit textual guidance, generated acoustic tokens become part of the context for subsequent text predictions. Given identical speech inputs, we observe markedly lower answer accuracy for the internal text generated in speech-to-text-and-speech (S2TS) mode than for speech-to-text (S2T) responses. We term this discrepancy the \emph{output-mode gap} (OMG). To reduce OMG, we propose \emph{Joint-Output On-Policy Distillation} (JO-OPD), which distills the model's stronger S2T policy into joint generation using student-generated S2TS trajectories. At each text position, the S2T teacher provides soft targets from a text-only projection of the student's preceding outputs, while the student predicts from the corresponding full interleaved history. A preservation objective further regularizes native non-text predictions. Experiments on Step-Audio-2-mini and Baichuan-Audio-Instruct reveal OMG across two interleaved generation architectures. On Step-Audio-2-mini, JO-OPD reduces OMG from 42.87 to 16.26 percentage points on Spoken-MQA and from 29.72 to 13.04 points on speech-rendered GSM8K, with little change in S2T accuracy and substantially larger reductions than matched SFT baselines. ASR-based evaluation further shows a 7.49-point improvement in spoken-answer accuracy on Spoken-MQA.