AI 中文总结
针对 Whisper 在无语音音频上产生幻觉的问题,提出同时抑制幻觉与恢复真实短语的微调方法,加入合成正样本,将静音幻觉从 61.9% 降至 2.4%,并提升短语恢复率。
AI 中文摘要
Whisper 仍然是在生产中运行的模型:一个宽松许可的检查点,支持 99 种语言,无需针对每种语言进行调优。但它会写出没有人说过的句子。在 42 段纯室内环境音片段上,whisper-large-v3 在其中 61.9% 的片段上输出了词语,并在 100% 的片段上输出了某些内容。通常的应对方法是先从 Whisper 的伪标签中蒸馏出一个学生模型,先从语料库中过滤掉幻觉,这会删除证据,同时保留行为:学生模型被拟合到教师模型正确的子集上,并继承了其自身训练数据中不存在的失败模式。第二种应对方法是添加非语音音频,让模型学会保持安静,这已经发布并有效,但只教会了抑制而没有区分。在所有非语音测试臂上表现最好的检查点也是恢复真正口语短语最少的检查点,并且删除了 58% 的重复语音。教师模型必须被修复,同时处理信号的两个部分。我们的基准测试同时评估这两方面:11,852 个片段分布在八个测试臂上,6,267 个合成正样本,8,296 个使生产模型失败的真实音频片段,以及 58 种语言的 FLEURS。我们收集了 100 种语言的 40,891 个幻觉短语,通过测量选择文本转语音系统,并将这些短语合成为正样本,使模型同时遇到相同的文本作为要抑制的内容和要转录的内容。在 33 对仅因这些正样本而不同的微调模型中,添加它们降低了 31 对中真实无语音音频上的词语输出,并在 32 对中提高了短语恢复率;在保持英语准确性的 30 对中,30 对全部有效。最佳检查点将静音上的幻觉从 61.9% 降至 2.4%,将真实无语音音频上的词语输出从 99.9% 降至 47.8%,同时将短语恢复率从 69.8% 提高到 82.7%。基准测试、词典和合成语料库已在此 https URL 发布。
英文摘要
Whisper is still what runs in production: one permissively licensed checkpoint, 99 languages, no per-language tuning. But it writes sentences nobody said. On 42 clips of pure room tone, whisper-large-v3 emits words on 61.9% of them and emits something on 100%. The usual response is to distil a student from Whisper pseudo-labels, filtering the hallucinations out of the corpus first, which deletes the evidence while leaving the behaviour in place: the student is fitted to the subset where the teacher was right, and inherits a failure mode absent from its own training data. The second response, adding non-speech audio so the model learns to stay quiet, already ships and works, but teaches suppression without discrimination. The checkpoint best on every non-speech arm here is also the one that recovers the fewest genuinely spoken phrases and deletes 58% of repeated speech. The teacher has to be fixed, with both halves of the signal. Our benchmark scores both at once: 11,852 clips over eight arms, 6,267 synthetic positives, 8,296 clips of real audio that made a production model fail, and FLEURS in 58 languages. We collect 40,891 hallucination phrases in 100 languages, choose a text-to-speech system by measurement, and synthesise those phrases as positives, so the model meets the same text both as something to suppress and as something to transcribe. Across 33 matched pairs of fine-tunes differing only by those positives, adding them lowers word emission on real voice-free audio in 31 pairs and raises phrase recovery in 32; among the 30 pairs that hold English accuracy it is 30 out of 30. The best checkpoint takes hallucination on silence from 61.9% to 2.4% and words over real voice-free audio from 99.9% to 47.8%, while raising phrase recovery from 69.8% to 82.7%. The benchmark, lexicon and synthetic corpus are released at https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination