Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center, School of Computer Science and Technology, Xinjiang University(新疆多模态智能处理与信息安全工程技术创新中心,计算机科学与技术学院,新疆大学) ; Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) ; School of Electrical Engineering and Automation, Tianjin University of Technology(电气工程与自动化学院,天津工业大学)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM
Comments Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2025
机构 * Santa Clara, CA, USA(美国圣克拉拉) ; Carnegie Mellon University(卡内基梅隆大学) ; Huazhong University of Science and Technology(华中科技大学) ; Nanjing University of Information Science and Technology(南京信息工程大学)
专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CL、eess.AS
Comments Accepted by Interspeech
Journal ref Proc. of Interspeech2025
机构 * Academia Sinica(台湾“中央研究院)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments Accepted to Interspeech 2025 Workshop
机构 * Tampere University, Tampere, Finland(塔尔库大学) ; University of Oxford, Oxford, UK(牛津大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Preprint version. The Version of Record is published in DAGM GCPR 2025 proceedings with Springer Lecture Notes in Computer Science (LNCS). Updated results and resources are available at the project page: https://saganet.notion.site
机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) ; Institute of Acoustics Chinese Academy of Science(中国科学院声学研究所)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments Submitted to ICASSP 2026
专题命中 音频语音多模态 :multimodal(abstract)
Comments 23 pages, 21 figures
机构 * Robotics Department, University of Michigan(密歇根大学机器人系)
专题命中 音频语音多模态 :multimodal(abstract)
Comments 8 pages, 7 figures