Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
探讨在音频深度伪造检测中使用多模态大语言模型的可行性
机构 * Indian Institute of Technology, Ropar, India(印度理工学院罗帕尔分校) ; Machine Intelligence Group, Birla Institute of Technology and Science, Pilani, Hyderabad Campus, India(比拉理工科学院帕利尼 Hyderabad 分校机器智能小组) ; Monash University, Melbourne, Australia(墨尔本大学)
专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV
AI总结 本文探讨了多模态大语言模型在音频深度伪造检测中的可行性,通过结合音频输入与多提示方法,展示了模型在域内数据上的良好表现及潜在应用价值。
Comments Accepted at IJCB 2025