SEAR:面向音频语言模型的伪造证据接地音频推理基准
SEAR: Spoofing Evidence-Grounded Audio Reasoning Benchmark for Audio Language Models
另 1 家 · 查看机构详情
- University of Surrey(萨里大学)
- Singapore University of Technology and Design(新加坡科技设计大学)
- Guangxi University(广西大学)
- Shenzhen University of Advanced Technology(深圳先进技术大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对现有音频深度伪造检测基准忽视声学证据验证的问题,提出SEAR基准和BAEA智能体,通过证据识别与量化等四项任务评估ALM,实验显示BAEA-Fixed提升判定与取证理由,且误导证据损害性能。
中文摘要 AI 辅助
音频语言模型(ALMs)越来越多地被用于音频深度伪造检测(ADD),然而现有基准仅评估其判定或理由的合理性,而未验证所依据的声学证据。为解决此问题,我们首先引入伪造证据接地音频推理(SEAR),这是一个四项任务的音频问答(AQA)基准,通过声学证据识别与量化、深度伪造检测及取证理由生成来评估基于ALM的ADD。我们进一步提出一种基于真实语音的声学证据智能体(BAEA),该智能体在固定或自适应证据获取策略下,为冻结的ALM配备受控的声学工具。使用六个ALM进行的实验揭示了合理理由与可验证的声学证据推理之间的明显差距,而BAEA-Fixed在两个评估分区上均提升了最终判定和取证理由。受控干预进一步表明,误导性证据会同时降低检测和接地性能。
英文摘要
Audio language models (ALMs) are increasingly used for audio deepfake detection (ADD), yet existing benchmarks assess their verdicts or rationale plausibility without verifying the underlying acoustic evidence. To address this issue, we first introduce spoofing evidence-grounded audio reasoning (SEAR), a four-task AQA benchmark to evaluate ALM-based ADD through acoustic evidence identification and quantification, deepfake detection, and forensic rationale generation. We further propose a bona-fide-based acoustic evidence agent (BAEA), which equips a frozen ALM with controlled acoustic tools under \textsc{fixed} or \textsc{adaptive} evidence-acquisition policies. Experiments with six ALMs reveal a clear gap between plausible rationales and verifiable acoustic evidence reasoning, while BAEA-\textsc{Fixed} improves final verdicts and forensic rationales on both evaluation partitions. Controlled interventions further show that misleading evidence degrades both detection and grounding performance.