发表机构
IISER Bhopal; King’s College London; MBZUAI(印度科学教育与研究学院博帕尔分校; 伦敦国王学院; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语音错误信息,提出VeriSpeak基准,发现文本到语音存在模态差距,检索加显式推理可提升语音事实核查准确率至86.1%。
AI 中文摘要
在线错误信息越来越多地以语音形式出现,例如新闻片段、播客、访谈、政治演讲和社交媒体视频,这催生了对能够直接从语音中验证声明的事实核查系统的需求。我们引入了VeriSpeak,一个用于研究大型音频语言模型(LALMs)中基于语音的事实验证的探针基准。VeriSpeak包含3,879条语音声明,涵盖时间、地理和关系事实,并具有平衡的真假标签。该基准旨在检验事实验证能力是否从文本迁移到语音,以及检索增强的LALMs能否利用文本证据正确支持或反驳语音声明。我们的实验揭示了一致的文本-语音模态差距:能够可靠验证书面声明的LALMs在相同声明以语音形式呈现时往往失败。此外,仅靠检索带来的收益有限,因为模型经常将检索到的证据与语音声明混淆。相比之下,检索结合显式推理改善了声明-证据比较,一个思维调优的LALM达到了86.1%的准确率。VeriSpeak强调,有效的语音错误信息检测不仅需要语音理解,还需要对检索到的证据进行基于推理的推理。该数据集通过Hugging Face公开提供,网址为https://this https URL。
英文摘要
Online misinformation increasingly appears in spoken formats such as news clips, podcasts, interviews, political speeches, and social media videos, creating a need for fact-checking systems that can verify claims directly from speech. We introduce VeriSpeak, a probe benchmark for studying speech-based fact verification in Large Audio Language Models (LALMs). VeriSpeak contains 3,879 spoken claims spanning temporal, geographical, and relational facts, with balanced true and false labels. The benchmark is designed to examine whether factual verification ability transfers from text to speech, and whether retrieval-augmented LALMs can use textual evidence to correctly support or refute spoken claims. Our experiments reveal a consistent text-speech modality gap: LALMs that verify written claims reliably often fail on the same claims when spoken. Moreover, retrieval alone provides limited gains because models frequently conflate retrieved evidence with the spoken claim. In contrast, retrieval combined with explicit reasoning improves claim-evidence comparison, with a thinking-tuned LALM reaching 86.1% accuracy. VeriSpeak highlights that effective speech misinformation detection requires not only speech understanding, but also grounded reasoning over retrieved evidence. The dataset is publicly available via Hugging Face at https://huggingface.co/datasets/abhiram4572/VeriSpeak.
CommentsAccepted to EMNLP (Main) 2026