AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
机构 * Nanjing University(南京大学)
专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.CV
Comments 21 pages, 11 figures
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Nanjing University(南京大学)
专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.CV
Comments 21 pages, 11 figures
机构 * Department of Computer Science and Engineering(计算机科学与工程系) ; The University of Texas at Arlington(德克萨斯理工大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS
Comments IEEE Transactions on Audio, Speech and Language Processing
机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学) ; Industrial Technology Research Institte of Shandong Province(山东省工业技术研究院)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS
Comments Published in ICASSPW 2024 (HSCMA)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS
Comments Accepted by 2022 IEEE International Conference on Big Data (IEEE Big Data 2022)
Journal ref IEEE BigData, Year: 2022; Pages: 3622-3630
机构 * Hefei University of Technology(合肥工业大学) ; Chinese Academy of Sciences(中国科学院) ; Beihang University(北航) ; Sangfor Technologies(深信服技术)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Accepted by CVPR 2025; Code: https://github.com/spyflying/VCT_AVS; Models: https://huggingface.co/nowherespyfly/VCT_AVS
机构 * MBZUAI ; Hefei University of Technology(合肥工业大学) ; University of Science and Technology of China(中国科学技术大学) ; National University of Singapore(新加坡国立大学) ; Peking University(北京大学) ; Tsinghua University(清华大学) ; OpenNLP Lab(OpenNLP实验室)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Technical Report
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM
Comments IEEE VIS 2025 Short Paper
专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS
Comments 14 pages, 10 figures
机构 * Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(深圳国际研究生院,清华大学,深圳,中国)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS
Comments Accepted by ICME2025
机构 * Ben Gurion University(本古里安大学) ; University of Michigan(密歇根大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL
Comments To be published in WOAH, July 2025. arXiv admin note: text overlap with arXiv:2409.14464
机构 * Hosei University(恒生大学)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Accepted to ICMR2025
机构 * School of Information and Communication Engineering(信息与通信工程学院)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Our paper has been accepted by ICME 2025
机构 * Indian Institute of Information Technology Kottayam(印度信息技术学院科塔亚姆) ; Indian Institute of Technology Jodhpur(印度理工学院朱罗普尔)
专题命中 音频语音多模态 :multi-modal(title);cross-modal(abstract);分类 cs.CV
Comments Accepted in Interspeech 2025
机构 * Sorbonne University(索邦大学) ; CNRS(国家科学研究中心)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL
机构 * Joint NTU-UBC Research Centre of Excellence in Active Living for the Elderly (LILY), NTU(联合NTU-UBC老龄化积极生活卓越研究中心(LILY),NTU) ; College of Computing and Data Science, Nanyang Technological University (NTU), Singapore(计算与数据科学学院,南洋理工大学(NTU),新加坡) ; Tan Tock Seng Hospital, Singapore(坦 tok sing 医院,新加坡) ; Woodlands Health, Singapore(伍德兰兹健康,新加坡)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL
Comments Accepted at Interspeech 2025
机构 * Alibaba Group(阿里巴巴集团) ; Singapore National University of Singapore(新加坡国立大学)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments Interspeech2025
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; Dartmouth College(达特茅斯学院)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV
Comments Project page: https://vincent2311.github.io/ReelWave_demo
机构 * NetEase Inc.(网易公司)
专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS
Comments ICMR 2025
机构 * Northeastern University(东北大学) ; University of Technology Sydney(悉尼大学) ; Xiamen University(厦门大学) ; Technical University of Munich(慕尼黑技术大学) ; University of Cambridge(剑桥大学) ; National Information Institute(国家信息研究所) ; Osaka University(大阪大学) ; Imperial College London(伦敦帝国理工学院)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
Comments This paper has been accepted as part of the MPDD Challenge in the ACMMM 2025 Grand Challenge
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院) ; Peking University(北京大学) ; School of AI for Science(科学人工智能学院) ; Peng Cheng Laboratory(鹏城实验室)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Accepted by IEEE TCSVT
机构 * Center for Machine Vision and Signal Analysis, University of Oulu(机器视觉与信号分析中心,奥卢大学) ; Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) ; State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
机构 * The University of Hong Kong(香港大学)
专题命中 音频语音多模态 :cross-modal(title,abstract);分类 eess.AS
Comments Accepted by Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (ACM IMWUT/UbiComp 2025)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM
机构 * Microsoft Research(微软研究院)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV
机构 * School of Computer Science and Technology(计算机科学与技术学院)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI
Comments Main paper (9 pages). Accepted for publication by ICMR(International Conference on Multimedia Retrieval) 2025
机构 * Sigmedia Group, School of Engineering Trinity College Dublin(Sigmedia集团,工程学院,三一学院都柏林)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments 5 pages, 2 figures. Accepted to ICASSP 2025