AI 中文总结
该研究定义SURE-Challenge评估语音大语言模型生成前的准入步骤,以Qwen2-Audio和能量加Whisper评分规则开展实验,发现了仅答案评分遗漏的生成前错误模式。
AI 中文摘要
语音大语言模型(Speech LLMs)通常在回答后才接受评估,而操作系统首先需要决定是否应将波形发送至该模型。我们为这一准入步骤定义了语音不支持拒绝评估挑战(Speech-Unsupported Rejection Evaluation Challenge,SURE-Challenge)。该基准将源自LibriSpeech的转录文本与首词问答任务,在不相交源划分下,搭配不支持的静音、有色噪声、合成音调和源模糊的 babble(多声源混响)。前端 ablation 实验使用 Qwen2-Audio;选定的能量加 Whisper 评分规则随后在6个语音/音频大语言模型上复现。在474行经泄漏筛选的 SURE-Extended 测试集上,原始 Qwen2-Audio 拒绝了204个不支持输入中的15个,而固定规则拒绝了204个中的196个,且未改变支持输入的准确率。外部检查限定了该数值:随着 Whisper 评分阈值收紧,Common Voice 的保留率下降;无速度 babble 在54个片段中,经不同随机种子处理后有18至24个被拒绝。该结果识别出仅基于答案评分所遗漏的生成前错误模式。
英文摘要
Speech language models (speech LLMs) can generate plausible outputs from audio that contains no usable speech evidence. We study this failure as a pre-generation support-estimation problem and present SURE-Voice, a training-free front end that decides whether an audio prompt contains intelligible speech evidence before calling a speech LLM. We build SURE-Challenge with a 640-example SURE-Core split and a 1,920-example SURE-Extended split derived from 120 LibriSpeech source utterances. Using one fixed operating point, an energy screen plus Whisper token confidence raises unsupported accuracy on the held-out Extended test from 0.000--0.133 to 0.919 for six non-degenerate speech LLM backbones, while supported accuracy remains 0.919--0.970 and downstream calls fall from 480 to 287. A 500-clip ESC-50 sanity set shows the same pattern on real environmental audio, with vocal non-speech as a residual failure mode. An overlap diagnostic shows that source attribution remains separate from speech-evidence filtering. The evidence supports a controlled benchmark baseline and a deployment-oriented analysis; it does not establish universal robustness to semantic answerability, gain variation or natural conversations.