arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迷失在语音中:跨音频与文本的三语语音幻觉检测

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

Meruyert Aristombayeva, Jason S. Lucas, Chaewan Chun, Dongwon Lee

arXiv 2608.24707首次发表:更新:

发表机构

Satbayev University; University of Colorado Boulder; The Pennsylvania State University(萨塔巴耶夫大学; 科罗拉多大学博尔德分校; 宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对低资源语言语音幻觉检测的空白,构建三语语音幻觉基准,评估多模态检测模型,发现文本检测更优,合成训练检测器在真实数据上迁移性强,还揭示合成基准的混淆因素。

AI 中文摘要

尽管基于文本的幻觉检测已得到广泛研究,但语音幻觉检测仍未被充分探索,尤其是针对低资源语言的情况。我们提出了首个多语言语音幻觉基准,包含12013个新闻样本,涵盖英语、俄语和哈萨克语,具有三种类型的可控幻觉及三个严重程度级别。样本包含原始文章以及文本和音频形式的对齐幻觉对应物。我们在合成语料库之外补充了290个经事实核查的假新闻条目,这些条目为俄语(225个)和哈萨克语(65个)原生收集,经翻译为其他语言并通过相同的TTS-ASR管道生成。我们评估了微调后的多语言编码器,以及在零样本上下文设置下基于文本与直接音频处理的多模态解码器模型。基于文本的检测通常优于直接音频处理,对于强编码器而言,二元任务性能下降与各语言ASR错误相关。在真实世界的假新闻上,经合成训练的检测器迁移性较强(原始文本上的宏F1值为0.82-0.88),而俄语来源分析揭示了与真实性相关及模型依赖的机器风格信号,量化了合成幻觉基准中的关键混淆因素。

英文摘要

While text-based hallucination detection is well studied, reference-free detection of factual alterations in speech remains underexplored, especially for low-resource languages. Our spoken benchmark comprises 12,013 English, Russian, and Kazakh news samples with three synthetic alteration types and three severity levels, pairing source articles with rewrites as text, synthesized audio, and ASR transcripts. We add 290 fact-checked misinformation items collected in Russian (225) and Kazakh (65), translated into the other language and rendered through the same TTS-ASR pipeline. We evaluate fine-tuned multilingual encoders and zero-shot multimodal decoders on text, transcripts, and audio. Detectors receive only target inputs without source articles or external evidence; the task evaluates reference-free classification rather than evidence-grounded verification. Encoder degradation from source text to transcripts generally tracks per-language ASR error on the binary task. Among decoders, only Gemma-3n exceeds the binary majority-class baseline in macro-F1, on transcripts only; the other four fall below their respective baselines. Comparisons between audio and transcripts are confounded by differences in evaluation coverage and class balance. Synthetic-trained detectors achieve 0.82--0.88 macro-F1 on real-world misinformation source text; Russian provenance analysis reveals veracity-related and model-dependent machine-style signals, a key confound in synthetic hallucination benchmarks.

CommentsAccepted to the SALMA (EMNLP workshop)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑