Real-Time System for Audio-Visual Target Speech Enhancement
机构 * Bose Corporation(博世公司) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments Accepted into WASPAA 2025 demo session
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Bose Corporation(博世公司) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments Accepted into WASPAA 2025 demo session
机构 * Wuhan University of Science and Technology(武汉科技大学)
专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 eess.AS
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV
Comments Conference on Robot Learning 2025
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments Accepted into Automatic Speech Recognition and Understanding- ASRU 2025
机构 * Johannes Kepler University Linz(约翰内斯·开普勒大学林茨) ; Human-centered AI Group, AI Lab, Linz Institute of Technology(以人为本的人工智能小组、人工智能实验室、林茨技术研究所)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM
Comments 7 pages, 6 tables, IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI)
机构 * Faculty of Engineering, Bar-Ilan University(巴伊兰大学工程学院) ; OriginAI
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
机构 * École Polytechnique Fédérale de Lausanne (EPFL), Switzerland(瑞士联邦理工学院洛桑校区) ; ETH Zürich, Switzerland(瑞士苏黎世联邦理工学院) ; T.H. Chan School of Public Health, Harvard University, USA(哈佛大学T.H. Chan公共卫生学院)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
Comments This paper was originally submitted to the CODEML workshop for ICML 2025. 9 pages (including references and appendices)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS
Comments Accepted at Large Language Models for Music & Audio Workshop (LLM4MA) 2025
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Tencent(腾讯)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
Comments Accepted by ACMMM 2025
机构 * Peraton Labs(珀顿实验室)
专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV
Comments 11 pages, 11 figures
机构 * Guangdong Provincial Key Laboratory of Quantum Engineering and Quantum Materials(广东省量子工程与量子材料重点实验室) ; School of Electronic Science and Engineering (School of Microelectronics), South China Normal University(华南师范大学电子科学学院(微电子学院)) ; North Carolina Central University(北卡罗来纳中央大学) ; Georgia Institute of Technology(佐治亚理工学院) ; Columbia University(哥伦比亚大学) ; LMU Munich & Munich Center for Machine Learning(慕尼黑大学及慕尼黑机器学习中心)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL
Comments EMNLP 2025 Findings
机构 * Inner Mongolia University(内蒙古大学) ; Center for Language and Speech Processing (CLSP)(语言与语音处理中心) ; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL
Comments Accepted by EMNLP 2025
机构 * University of Science and Technology of China(科学技术大学) ; AI Research Center, Midea Group (Shanghai) Co.,Ltd.(美的集团(上海)有限公司人工智能研究中心)
专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS
Comments Accepted by ASRU 2025
机构 * Campus Fryslân, University of Groningen(格罗宁根大学弗里桑校区)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL
Comments 20 pages, 7 figures, Submitted to IEEE Transactions on Affective Computing
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM
Comments Accepted for publication at the 26th International Society for Music Information Retrieval Conference (ISMIR 2025)
机构 * Google(谷歌)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
Comments Accepted in IEEE COMPSAC 2025
Journal ref 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC)
机构 * Shanghai Jiao Tong University(上海交通大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM
机构 * Department of Computer Science and Engineering, Koç University(计算机科学与工程系,科克大学) ; Department of Psychology, Boğaziçi University(心理学系,博多伊大学) ; National Institute of Advanced Industrial Science and Technology (AIST), Intelligent Platforms Research Institute(国家先进工业科学与技术研究院(AIST),智能平台研究机构) ; Department of Psychology, Boğaziçi University University(心理学系,博多伊大学) ; Department of Computer Engineering, Hacettepe University(计算机工程系,哈切塞特佩大学) ; KUIS AI Center(KUIS人工智能中心)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Accepted for publication in IEEE Transaction on Pattern Analysis and Machine Intelligence (IEEE TPAMI)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments Preprint of the paper presented at Euronoise 2025 Malaga, Spain
机构 * Human-AI Interaction (HAIx) Lab, IIT Gandhinagar(人机交互(HAIx)实验室,印度加尔各答理工学院)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
机构 * Ritsumeikan University(立命馆大学) ; Soka University(早稻田大学) ; Kyoto University(京都大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
Comments See website at https://emergentsystemlabstudent.github.io/MIEL/. Accepted at IEEE RO-MAN 2025
机构 * Ziane Achour University(赞赞·阿赫尔大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL
机构 * Beijing Fosafer Information Technology Co., Ltd.(北京福萨弗信息科技有限公司)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS
Comments 2 pages. pubilshed by ICASSP2025
机构 * Institute of Information Engineering(信息工程研究所) ; Chinese Academy of Sciences(中国科学院) ; School of Cyber Security University of Chinese Academy of Sciences(中国科学院网络安全学院) ; Deakin University(德肯大学) ; Harbin Engineering University(哈尔滨工程大学) ; Institute of Computing Technology Chinese Academy of Sciences(中国科学院计算技术研究所) ; Institute of Forensic Science Ministry of Public Security(公安部刑事科学技术研究所)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Accepted by NeurIPS 2024
Journal ref Advances in Neural Information Processing Systems, Volume 37, Pages 86124-86144, Year 2024
机构 * Institute of Artificial Intelligence (TeleAl), China Telecom(人工智能研究院(TeleAl),中国电信) ; School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学)
专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 eess.AS
机构 * VNU University of Engineering and Technology(越南工程大学) ; Delft University of Technology(代尔夫特理工大学)
专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI
机构 * East China Normal University(华东师范大学) ; Shanghai Jiao Tong University(上海交通大学) ; City University of Hong Kong(香港城市大学) ; Shanghai University of Electric Power(上海电力大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV
Comments The proposed method achieves first place in the ICCV VQualA 2025 EVQA-SnapUGC Challenge on short-form video engagement prediction
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments 9 pages
机构 * School of Music(音乐学院) ; Research(研究)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments Accepted into Interspeech 2025; corrected author name typo
机构 * Ben Gurion University of the Negev(本· Gurion 内盖夫大学)
专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI