arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2025-12-04 至 2025-12-04 共收录 1 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 音视频/视觉语言融合 1 篇

2508.17282 2025-12-04 cs.AI cs.SD 50%

ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection

ERF-BA-TFD+: 一种用于音频视觉深度伪造检测的多模态模型

Xin Zhang, Jiaming Chu, Jian Zhao, Yuchu Jiang, Xu Yang, Lei Jin, Chi Zhang, Xuelong Li

专题命中 音视频/视觉语言融合 :audio-visual fusion(abstract)

AI总结 ERF-BA-TFD+通过结合增强接收场和音频视觉融合,提出了一种多模态深度伪造检测模型,在DDL-AV数据集上实现了最先进的检测性能。

Comments The paper is withdrawn after discovering a flaw in the theoretical derivation presented in Section Method. The incorrect step leads to conclusions that are not supported by the corrected derivation. We plan to reconstruct the argument and will release an updated version once the issue is fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏