arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

倾听,勿抄袭:内化音频基础的支架上下文以实现稳健的全模型语音理解

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

Pengfei Zhang, Biao Tian, Tianxin Xie, Minghao Yang, Xiangang Li, Li Liu

arXiv 2607.21943首次发表:更新:

AI 中文总结

研究针对全模型在嘈杂重叠语音中准确率下降问题,提出音频基础的支架上下文(AGSC)方法,通过三步构建线索,经测试优化后用于训练,降低了无线索错误率,还制定联合任务,内化后几乎不增加推理开销。

AI 中文摘要

全模型在转录清晰的单说话者语音时表现良好,但在说话者重叠且场景嘈杂时,其准确率会大幅下降,而此时知道谁说了什么最为重要。自然的解决方法是简短的场景描述。我们展示了这种方法为何有风险:带答案的文本会让模型抄袭而非倾听,导致分数上升但实际未听到任何内容;无声测试能立刻暴露这种捷径。我们将这种失败模式称为感知绕过,并通过音频基础的支架上下文(AGSC)来解决。AGSC 连接三个步骤:首先,从音频构建线索以引导倾听而不给出答案;其次,通过答案重叠和无声测试探测线索是否有泄漏及对音频的依赖;最后,这些线索用于训练但在测试时消失,产生无线索能力。在三个异构全模型上,基于 AGSC 训练可将重叠嘈杂语音上的无线索上限平均排列词错误率(mpWER)从 25%-71%降至 9%-15%。对于流控制,我们制定了联合 GDPO 任务,模型在其中学习何时使用线索以及如何从单独归一化的格式、门控和转录奖励中生成说话者归因的转录本。内化后,AGSC 几乎不增加推理开销。

英文摘要

Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text lets the model copy instead of listen, so the score rises although nothing has been heard; a silent test exposes this shortcut at once. We call this failure mode perception bypass and address it with Audio-Grounded Scaffold Context (AGSC). AGSC links three steps: first, we build clues from audio to guide listening without giving the answer; second, answer-overlap and silence tests probe them for leakage and audio dependence; finally, those clues scaffold training but vanish at test time, yielding no-clue capability. Across three heterogeneous Omni models, training on AGSC lowers no-clue capped mean permutation word error rate (mpWER) on overlapping, noisy speech from 25%-71% to 9%-15%. For streaming control, we formulate a joint GDPO task in which the model learns when to use a clue and how to produce a speaker-attributed transcript from separately normalized format, gate, and transcript rewards. After internalization, AGSC adds almost no inference overhead.

Comments9 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑