ThinkOmni:用于音频伪造检测与定位的推理驱动全模态大语言模型框架
ThinkOmni: A Reasoning-Driven Omni-Modal LLM Framework for Audio Forgery Detection and Localization
浏览论文内容
中文总结 AI 辅助
针对现有AFDL方法泛化能力不足的问题,提出推理驱动的全模态大语言模型ThinkOmni,构建FACoT数据集,引入FMIL与FCML,实现了跨数据集的音频伪造检测与定位。
中文摘要 AI 辅助
现有音频伪造检测与定位(AFDL)方法常过拟合数据集特有的低级伪影,限制了其对细微、局部及未见过的操纵的泛化能力。近期基于音频大语言模型(ALLM)的方法将AFDL视为问答任务,但仍隐式建模取证证据,未将操纵线索与预测关联。为弥合此差距,我们提出ThinkOmni,这是一个推理驱动的全模态大语言模型,可联合执行显式取证推理、伪造检测和时序操纵定位。为实现显式推理监督,我们构建了Forensic-Aware Chain-of-Thought(FACoT),这是一个包含结构化取证证据和推理标注的10万样本数据集。利用FACoT,我们引入了Forensic-Aware Modality-Incremental Learning(FMIL),它逐步将语义、声学和频谱-视觉表示与LLM主干对齐,以捕捉互补的取证线索。我们进一步提出Forensic-Consistent Multi-task Loss(FCML),它结合加权交叉熵与自适应定位损失,以协调推理生成、伪造检测和时序定位。大量实验表明,ThinkOmni在检测和定位任务中均实现了强大的跨数据集泛化能力。代码、模型、数据和推理示例可在此https URL获取。
英文摘要
Existing audio forgery detection and localization (AFDL) methods often overfit dataset-specific low-level artifacts, limiting their generalization to subtle, localized, and unseen manipulations. Recent audio large language model (ALLM)-based approaches cast AFDL as question answering but still model forensic evidence implicitly, without linking manipulation cues to predictions. To bridge this gap, we propose ThinkOmni, a reasoning-driven omni-modal large language model that jointly performs explicit forensic reasoning, spoofing detection, and temporal manipulation localization. To enable explicit reasoning supervision, we construct Forensic-Aware Chain-of-Thought (FACoT), a 100K-sample dataset with structured forensic evidence and reasoning annotations. Leveraging FACoT, we introduce Forensic-Aware Modality-Incremental Learning (FMIL), which progressively aligns semantic, acoustic, and spectral-visual representations with the LLM backbone to capture complementary forensic cues. We further propose Forensic-Consistent Multi-task Loss (FCML), which combines weighted cross-entropy with an adaptive localization loss to coordinate reasoning generation, spoofing detection, and temporal localization. Extensive experiments show that ThinkOmni achieves strong cross-dataset generalization in both detection and localization. Code, models, data, and inference examples are available at https://beyond0814.github.io/ThinkOmni/.
发表机构
- Shenzhen University(深圳大学)
- Afirstsoft Technology Group Co., Ltd.(安福软件科技集团有限公司)
机构由 AI 辅助整理,请以论文原文为准。