arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越语言和模态限制学习说话者身份:来自POLY-SIM 2026挑战赛的见解

Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge

Marta Moscati, Muhammad Saad Saeed, Marina Zanoni, Mubashir Noman, Rohan Kumar Das, Monorama Swain, Yassin Terraf, Yufang Hou, Elisabeth Andre, Khalid Mahmood Malik, Markus Schedl, Shah Nawaz

arXiv 2607.13669首次发表:更新:

发表机构

Institute of Computational Perception, Johannes Kepler University; University of Michigan-Flint; BDO DIGITAL Gmbh; Mohamed bin Zayed University of Artificial Intelligence; Fortemedia; College of Computing, Mohammed VI Polytechnic University; IT:U Interdisciplinary Transformation University; University of Augsburg(约翰内斯·开普勒大学计算感知研究所; 密歇根大学弗林特分校; BDO数字有限公司; 穆罕默德·本·扎耶德人工智能大学; 富特媒体公司; 穆罕默德六世理工大学计算学院; IT:U跨学科转型大学; 奥格斯堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

POLY-SIM 2026挑战赛旨在解决多模态说话者识别中语言和模态限制问题,针对实际应用中信息缺失、多语言等挑战,为比较解决方案提供标准化设置。

AI 中文摘要

多模态说话者识别系统通常假设在训练和测试期间完整且同质的视听模态可用,且每个说话者只说一种语言。但在实际应用中这些假设往往不成立,视觉或音频信息可能缺失,多语言说话者也带来复杂性。POLY-SIM 2026挑战赛旨在解决这些问题,并为比较所提解决方案提供标准化设置。

英文摘要

Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing, and assume each speaker only speaks a single language. However, in real-world applications, such assumptions often do not hold. Visual or audio information may be missing due to occlusions, camera or microphone failures, or privacy constraints. Multilingual speakers introduce additional complexity due to linguistic variability across languages. These situations constitute substantial challenges for the robustness and generalization capabilities of multimodal speaker identification systems. Aim of the POLY-SIM 2026 challenge is to address these aspects of speaker identification and to provide a standardized setup for the comparison of the proposed solutions.

CommentsAccepted at ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑