arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03620eess.AScs.AIcs.SD

ToolDF:面向混合真实性音频深度伪造检测的工具集成推理框架

ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection

Taewoo Kim, Young Han Lee, Nam In Park, Chanwoo Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对混合真实性音频深度伪造检测问题,提出ToolDF工具集成推理框架,引入混合真实性ADD基准,在复合检测任务上性能优于基线,同时提供可解释证据。

中文摘要 AI 辅助

音频深度伪造检测通常被建模为单域音频的片段级二分类任务。然而,现实中经过操纵的音频可能呈现混合真实性,即真实与操纵线索在时间过渡、重叠声源或两者兼而有之的情况下共存。该场景不仅需要检测操纵音频,还需要定位为决策提供证据的组件。我们提出ToolDF,这是一个用于混合真实性音频深度伪造检测的工具集成推理框架。ToolDF采用音频大语言模型作为编排器,通过监督工具使用轨迹进行训练,它能自适应分析音频场景,选择性执行声源分离,将组件路由到特定领域的专家,并将其证据聚合成可解释的裁决。我们还引入了混合真实性音频深度伪造检测(ADD)基准,涵盖时间过渡、声学重叠和混合混合情况。实验结果表明,ToolDF在复合类型检测上取得了最佳整体性能,相比最强的整体基线和固定流水线,其宏F1值分别提升了3.72和14.39个百分点,同时提供了定位到时间区域和声源的可解释证据。我们的源代码和数据集在网上公开可用。

英文摘要

Audio deepfake detection is commonly formulated as clip-level binary classification of single-domain audio. However, real-world manipulated audio can exhibit mixed authenticity, where genuine and manipulated cues coexist across temporal transitions, overlapping sources, or both. This setting requires not only detecting manipulated audio but also localizing the components that provide evidence for the decision. We propose ToolDF, a tool-integrated reasoning framework for mixed-authenticity audio deepfake detection. ToolDF employs an audio large language model as an orchestrator trained with supervised tool-use trajectories. It adaptively analyzes the audio scene, selectively performs source separation, routes components to domain-specific experts, and aggregates their evidence into an interpretable verdict. We further introduce a mixed-authenticity ADD benchmark covering temporal transitions, acoustic overlaps, and hybrid mixtures. Experimental results show that ToolDF achieves the best overall performance on composite-type detection, achieving macro-F1 gains of 3.72 and 14.39 points over the strongest monolithic baseline and a fixed pipeline, respectively, while providing interpretable evidence localized to temporal regions and acoustic sources. Our source code and dataset are publicly available online.

发表机构

  • Multi-Modal Research Center, KETI(韩国电子通信研究院多模态研究中心)
  • Korea University(高丽大学)
  • Digital Analysis Section, National Forensic Service(国家法医服务局数字分析科)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑