SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM
机构 * Google(谷歌)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV
Journal ref Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST '25), Article 17, 1-16, 2025
专题命中 音频语音多模态 :multimodal(title,abstract)
机构 * University of Pennsylvania(宾夕法尼亚大学)
专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL
机构 * South China Normal University, Guangzhou, China(华南师范大学) ; Xiamen Rekey Medical Technology Co., LTD, Xiamen, China(厦门瑞康医疗科技有限公司)
专题命中 音频语音多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI
Comments 8 pages, 2 figures. To appear in: Proceedings of the 28th European Conference on Artificial Intelligence (ECAI 2025), Frontiers in Artificial Intelligence and Applications, Vol. 413. DOI: 10.3233/FAIA251182
机构 * University of Tsukuba(茨口大学) ; Cluster Metaverse Lab(元宇宙集群实验室) ; The University of Tokyo(东京大学) ; ZOZO Research(ZOZO研究所)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.MM
Comments 3 pages, 2 figures, 1 table. Presented at SIGGRAPH Asia 2025 Posters (SA Posters '25), December 15-18, 2025, Hong Kong, Hong Kong
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; Prometheus Vision Technology Co., Ltd.(普罗米修斯视觉科技有限公司) ; The Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV
Comments Proceedings of the Computer Graphics International 2025 (CGI'25)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL
Comments This article will serve as an extension of the preceding work, "VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models" (arXiv:2505.15727). Therefore, we have chosen to withdraw to avoid potential duplicate publication. We will update the previously open-sourced paper of VocalBench in several weeks to include the content of VocalBench-zh
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments Our paper has been accepted to IEEE TPAMI