SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);MLLM(abstract);cross-modal(abstract)
Comments Project page: https://github.com/DianJin-HFUT/SimToken
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);MLLM(abstract);cross-modal(abstract)
Comments Project page: https://github.com/DianJin-HFUT/SimToken
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学) ; Zhejiang University - Ant Group Joint Lab of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室) ; National University of Singapore, NUS-NCS Joint Lab(新加坡国立大学NUS-NCS联合实验室)
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM
Comments EMNLP 2025 Main Conference. 23 pages (8+ for main); 25 figures; 1 table
机构 * 1Laboratory of Cognitive Computing ; Application, College of Intelligence ; Computing, Tianjin University, Tianjin, China 2School of Computer Science, Shanghai Jiao Tong University, Shanghai, China 3School of Electrical \& Electronic Engineering, Nanyang Technological University, Singapore 4Huiyan Technology (Tianjin) Co., Ltd, Tianjin, China
专题命中 音频语音多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CL、cs.MM、eess.AS
Comments Submitted to ICASSP 2026
机构 * Trans-disciplinary Bachelor Degree Program National Taiwan University(台湾国立大学跨学科学士学位计划)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI
机构 * Wuhan University of Science and Technology(武汉科技大学)
专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 eess.AS
专题命中 音频语音多模态 :multimodal(title,abstract)
Comments Accepted by IEEE Transactions on Audio, Speech and Language Processing
机构 * Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室) ; Shenzhen Institute of Advanced Technology(深圳先进技术研究院) ; Chinese Academy of Sciences(中国科学院) ; University of Chinese Academy of Sciences(中国科学院大学) ; Tongyi Laboratory(通义实验室) ; University of New South Wales(新南威尔士大学) ; National University of Singapore(新加坡国立大学) ; University of Science and Technology of China(中国科学技术大学) ; MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition(脑启发智能感知与认知重点实验室)
专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL