Shared Multi-modal Embedding Space for Face-Voice Association
用于人脸-语音关联的共享多模态嵌入空间
专题命中 音频语音多模态 :multi-modal(title);分类 cs.CV
AI总结 本文提出了一种基于共享多模态嵌入空间的人脸-语音关联方法,通过自适应角边距损失实现跨语言的高效识别
Comments Ranked 1st in Fame 2026 Challenge, ICASSP
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
用于人脸-语音关联的共享多模态嵌入空间
专题命中 音频语音多模态 :multi-modal(title);分类 cs.CV
AI总结 本文提出了一种基于共享多模态嵌入空间的人脸-语音关联方法,通过自适应角边距损失实现跨语言的高效识别
Comments Ranked 1st in Fame 2026 Challenge, ICASSP
机构 * Adobe Research(Adobe研究院) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; MIT(麻省理工学院)
专题命中 音频语音多模态 :multimodal(title);分类 eess.AS
Comments Submitted to ICASSP 2026
机构 * Big Data Mining and Analytics, xxxxxxx 20xx, x(x): xxx-xxx(大数据挖掘与分析,xxxxx 20xx, x(x): xxx-xxx)
专题命中 音频语音多模态 :multimodal(title);分类 cs.CV
Comments Accepted for publication in Big Data Mining and Analytics (BDMA), 2025
专题命中 音频语音多模态 :multimodal(title);分类 eess.AS
Journal ref Interspeech 2024
专题命中 音频语音多模态 :multi-modal(title);分类 eess.AS
Comments 5 pages, 2 figures, 1 table. Accepted for presentation at Interspeech 2025
专题命中 音频语音多模态 :multi-modal(title);分类 eess.AS
Comments Accepted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
Journal ref Proceedings of the 25th International Society for Music Information Retrieval Conference, 705-712. San Francisco, California, USA and Online, November 10-14, 2024
机构 * School of Automation Science and Engineering, Xi’an Jiaotong University, Xi’an, China(自动化科学与工程学院,西安交通大学,西安,中国) ; School of Software Engineering, Xi’an Jiaotong University, Xi’an, China(软件工程学院,西安交通大学,西安,中国)
专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV
机构 * Department of Computer Science, BRAC University(计算机科学系,布拉克大学)
专题命中 音频语音多模态 :multi-modal(title);分类 cs.CV
Comments 11 pages, 12 figures
机构 * SimulaMet
专题命中 音频语音多模态 :multimodal(title);分类 cs.CV
Comments 20 pages, 9 figures, 9 tables
机构 * University of Southern California(南加州大学) ; Boston University(波士顿大学)
专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS
Comments In review for ICASSP 2024, 5 pages
机构 * Federal University of Goiás(戈亚斯联邦大学) ; Federal University of Rio Grande do Norte(北里奥格兰德联邦大学) ; Aeronautics Institute of Technology(航空技术研究所) ; Federal University of Mato Grosso(马托格罗索联邦大学)
专题命中 音频语音多模态 :multimodal(title);分类 cs.CL
机构 * LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI,巴黎电信学院,巴黎高等理工学院)
专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS
机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) ; Beijing Institute of Technology(北京理工大学)
专题命中 音频语音多模态 :multimodal(title);分类 cs.AI
Comments 7 pages, 5 figures, CHI2025
专题命中 音频语音多模态 :multi-modal(title);分类 eess.AS
Comments Published in: IEEE/ACM Transactions on Audio, Speech, and Language Processing ( Volume: 32), DOI: 10.1109/TASLP.2024.3497586
专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV
Comments Project Page: https://sunyasheng.github.io/projects/COSH-DIT
专题命中 音频语音多模态 :multimodal(title);分类 cs.CL
专题命中 音频语音多模态 :multimodal(title);分类 eess.AS
专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV
专题命中 音频语音多模态 :multimodal(title);分类 cs.CV
Comments Accepted (Poster) - NeurIPS 2024 Workshop MusIML
专题命中 音频语音多模态 :multi-modal(title);分类 eess.AS
Comments Accepted at INTERSPEECH 2024
专题命中 音频语音多模态 :any-to-any(title);分类 eess.AS
Comments Accepted at INTERSPEECH 2024
专题命中 音频语音多模态 :multimodal(title);分类 cs.CL
专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS
Comments INTERSPEECH 2024
专题命中 音频语音多模态 :multimodal(title);分类 cs.CV
专题命中 音频语音多模态 :multi-modal(title);分类 cs.MM
Comments 14 pages, In the 4th International Conference on SMART MULTIMEDIA, 2024
专题命中 音频语音多模态 :multimodal(title);分类 eess.AS
Comments Master's thesis
专题命中 音频语音多模态 :multimodal(title);分类 cs.CL
Comments This paper is part of the proceedings of the Dialogue Robot Competition 2023
专题命中 音频语音多模态 :multimodal(title);分类 cs.AI
专题命中 音频语音多模态 :multimodal(title);分类 eess.AS
Comments 5 pages
专题命中 音频语音多模态 :cross-modal(title);分类 eess.AS
Comments Accepted by InterSpeech 2022