AI 中文总结
本文提出结合跨模态注意力机制的凸特征嵌入方法,解决人脸与声音跨模态关联中的特征异质性问题,在VoxCeleb数据集的相关任务上较现有方法有显著提升。
AI 中文摘要
人脸与声音的关联学习是深度学习中最具挑战性的任务之一。本文提出一种简单但有效的跨模态特征嵌入方法,用于人脸与声音的关联。此前研究已针对跨模态关联任务建立声音片段与人脸图像的相关性,这些工作解决了跨模态区分问题,但低估了处理音频与视频间模态间特征异质性的重要性,导致大量假正例和假负例。为解决该问题,所提方法通过在跨模态特征间引入另一特征,学习跨模态特征的嵌入,促进同一人的声音和人脸特征嵌入到凸包中。此外,将跨模态注意力机制与凸嵌入技术结合,是减少假正例和假负例的高效策略,该策略通过最小化类间差异实现。我们在大规模VoxCeleb数据集上,对所提方法的跨模态验证、匹配和检索任务进行了详尽评估,大量实验结果表明,所提方法较现有最先进方法取得了显著提升。
英文摘要
Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal feature embedding method for the association of faces and voices. Previous work has studied cross-modal association tasks to establish the correlation between voice clips and facial images. These works have addressed cross-modal discrimination but underestimate the importance of handling heterogeneity in inter-modal features between audio and video, resulting in a lot of false positives and false negatives. To tackle the problem, the proposed method learns the embeddings of cross-modal features by making another feature exist between cross-modal features, facilitating the voice and face features of the same person to be embedded in a convex hull. Moreover, the incorporation of cross-modal attention mechanisms with convex embedding techniques represents a highly effective strategy for the attenuation of false positives and false negatives, accomplished via the minimization of inter-class discrepancies. We exhaustively evaluated our method for cross-modal verification, matching, and retrieval tasks on the large-scale VoxCeleb dataset. Extensive experimental results demonstrate that the proposed method achieves notable improvements over existing state-of-the-art methods.
Journal refMultimedia Systems, vol. 31, no. 4, pp. 296, July 2025
DOI:10.1007/s00530-025-01872-9