arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17913cs.CV

跨语言与跨性别的人脸-语音关联(FLAG)2027挑战赛评估计划

Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge Evaluation Plan

Marta Moscati, Swapnil Khandoker, Muhammad Saad Saeed, Shah Nawaz, Fatima Noor, Rohan Kumar Das, Mubashir Noman, Junaid Mir, Muhammad Haroon Yousaf, Khalid Malik, Markus Schedl

首次发表
浏览论文内容

中文总结 AI 辅助

针对人脸-语音关联模型依赖语言和性别线索的问题,提出FLAG 2027挑战赛,通过跨模态验证任务和两种评估设置,推动开发捕捉身份特定特征的模型。

中文摘要 AI 辅助

人脸-语音关联模型可能依赖于语音中的语言或性别线索,而非说话者特定的语音特征,这可能导致模型在识别多语言说话者或区分同性别说话者时性能下降。为了研究这些问题,我们提出了跨语言与跨性别的人脸-语音关联(FLAG)2027挑战赛。该挑战将人脸-语音关联定义为一个跨模态验证任务:给定一段语音,从由说话者人脸和一组负样本组成的“图库”中识别出说话者的人脸。模型在训练数据中未出现的身份(“未见”)上进行评估,并且同时针对训练数据中出现或未出现的语言(“已听”和“未听”)进行评估。采用两种评估设置来测试模型对性别的依赖:一种标准的、无约束的设置和一种性别约束的设置,后者使用同性别图库。现有基线模型在这些设置下的表现表明,模型在语言转换和性别约束设置下性能下降,凸显了需要开发能够捕捉超越语言和性别的身份特定方面的模型。该挑战提供了一个基准数据集、预训练基线模型和一个评估框架,以推动人脸-语音关联领域的发展。

英文摘要

Face--voice association models may rely on language or gender cues in the voice rather than on speaker-specific voice characteristics, which can lead to a performance deterioration when the model has to identify a multilingual speaker or distinguis same-gender speakers. To investigate these issues, we introduce the Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge. The challenge formulates face--voice association as a cross-modal verification task: given a voice, identify the speaker's face from a ``gallery'' of faces consisting of the speaker's face and a set of negative samples. Models are evaluated on identities not present in the training data (``unseen'') and both for languages present or absent from the training data (``heard'' and ``unheard''). Two evaluation settings are used to test models' reliance on gender: a standard, unconstrained and a gender-constrained one, where the latter uses a same-gender gallery. The performance of existing, baseline models in these settings reveals that models performance degrades under language shifts and in gender-constrained settings, highlighting the need to foster the development of models that capture identity-specific aspects beyond language and gender. The challenge provides a benchmark dataset, pretrained baseline models, and an evaluation framework to advance face--voice association.

发表机构

  • Johannes Kepler University(约翰·开普勒大学)
  • University of Michigan(密歇根大学)
  • University of Engineering and Technology(工程技术大学)
  • Fortemedia Singapore(Fortemedia新加坡公司)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Linz Institute of Technology(林茨理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑