Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM
Comments 16 pages, 9 figures, AAAI 2026
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM
Comments 16 pages, 9 figures, AAAI 2026
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.MM
Journal ref C. Li, Q. Yan, M. Kim, Z. Li, Y. Xu and L. -F. Yu, "Crafting Dynamic Virtual Activities with Advanced Multimodal Models," 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 120-130
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV
机构 * Hong Kong University of Science and Technology(香港科技大学) ; Kyoto University(京都大学) ; Georgia Institute of Technology(佐治亚理工学院) ; The University of Texas at Dallas(德克萨斯大学达拉斯分校) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; Cornell University(康奈尔大学) ; Indiana University(印第安纳大学) ; National Taiwan Normal University(台湾师范大学) ; University of Liverpool(利物浦大学) ; University of Edinburgh(爱丁堡大学) ; Zhejiang University(浙江大学) ; Purdue University(普渡大学) ; Emory University(埃默里大学) ; West China Biomedical Big Data Center, West China Hospital, Sichuan University(西京生物大数据中心,西京医院,四川大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI
Comments 25 pages, 5 tables
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted to AAAI-2026 Oral
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
Comments 7 pages, 5 figures
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.MM
Comments 8 pages, 10 figures
机构 * School of Electrical Engineering and Computer Science (EECS), KTH Royal Institute of Technology(电气工程与计算机科学学院(EECS),皇家理工学院) ; Department of Computer and Network Engineering, The University of Electro-Communications(计算机与网络工程系,东京电讯大学)
专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.MM
机构 * University at Buffalo, SUNY(布法罗大学,SUNY) ; Dolby Laboratories Inc.(杜比实验室公司)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM
机构 * College of Computer Science, Inner Mongolia University, China(内蒙古大学计算机科学学院) ; Lenovo, China(联想公司)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
专题命中 音频语音多模态 :audio-visual(title,abstract)
机构 * MERaLiON Team Institute for Infocomm Research (I 2 R), A*STAR, Singapore(MERaLiON团队信息与通信研究所(I 2 R),A*STAR,新加坡)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
机构 * Nanyang Technological University (NTU)(南洋理工大学) ; MiroMind(米罗Mind) ; Institute for Infocomm Research (I 2 R)(信息与通信研究院) ; A*STAR(科技研究局)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments Link: https://github.com/AudioLLMs/AudioBench/tree/main/IFEval-Audio
机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) ; DEVCOM Army Research Laboratory(国防部陆军研究实验室)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
专题命中 视频多模态 :cross-modal(title,abstract)
专题命中 视频多模态 :cross-modal(title,abstract)
Comments 12 pages, 13 figures, 3 tables, 2 algorithms
机构 * Anhui Provincial International Joint Research center for Advanced technology in Medical imaging(安徽省国际联合先进医学影像技术研究中心) ; School of Artificial Intelligence(人工智能学院) ; Anhui University(安徽大学) ; State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) ; Anhui Provincial Key Laboratory of Secure Artificial Intelligence(安徽省安全人工智能重点实验室) ; Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(粤港澳大湾区人工智能与数字经济实验室(深圳))
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
机构 * Department of Information Science at Cornell University(康奈尔大学信息科学系) ; Weill Cornell Medicine, Cornell University(韦尔医学院,康奈尔大学)
专题命中 视频多模态 :multimodal(abstract)
Comments This is the author's original submitted version of the paper accepted to the 2025 IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). \c{opyright} 2025 IEEE. Personal use of this material is permitted. For any other use, please contact IEEE
Journal ref 2025 34th IEEE International Conference on Robot and Human Interactive Communication RO-MAN pp. 1460-1465
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)
机构 * CMU(卡内基梅隆大学) ; The University of Tokyo(东京大学) ; The University of Toronto(多伦多大学)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL
机构 * Faculty Of Information Technology, VNU University of Engineering and Technology(信息科技学院,越南工程与技术大学) ; IT-BT Convergence Technology Division, Vietnam-Korea Institute of Science and Technology(IT-BT融合技术部,越南-韩国科学技术院) ; TADI Global Lab, TADI Global Company Limited(TADI全球实验室,TADI全球公司) ; Faculty of Finance, Banking Academy of Vietnam(金融学院,越南银行学院)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments 7 pages, 3 figures, 3 tables, FAIR 2025 conference
专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI
Comments AAAI 2026
机构 * School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) ; Key Laboratory of Sustainable Tourism Smart Assessment Technology, Ministry of Culture and Tourism, Sun Yat-sen University(文化旅游可持续评估技术重点实验室,中华人民共和国文化和旅游部,中山大学) ; Beijing Normal University(北京师范大学) ; Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL
Comments CCKS 2025 Shared Task Paper
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、eess.AS
机构 * Hong Kong Baptist University(香港 Baptist 大学) ; Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) ; National University of Singapore(新加坡国立大学) ; Beijing Normal University(北京师范大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments 28 pages, 14 figures, 19 tables
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 8 pages, 11 tables and figures
机构 * Kalinga Institute of Industrial Technology (KIIT)(喀里亚理工学院) ; Indian Institute of Technology (IIT), Bhubaneswar(印度理工学院(班加罗尔))
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL
Comments Accepted to IJCNLP-AACL Findings 2025
专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV
Comments Accepted by AAAI-2026
机构 * FAIR at Meta(Meta 的 FAIR 研究组) ; Meta Reality Labs(Meta 现实实验室) ; University of Southern California(南加州大学)
专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL、cs.AI
Comments Website: https://facebookresearch.github.io/DigiData
机构 * University of Texas at El Paso(德克萨斯理工大学) ; Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校)
专题命中 多模态评测 :cross-modal(abstract);分类 cs.AI