arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-13 至 2025-11-13 共收录 45 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2511.08971 2025-11-13 cs.HC cs.CV cs.MM 84%

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

Sicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun, You He, Jiankang Deng, Hang Zhang, Jifei Song, Zhensong Zhang

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 16 pages, 9 figures, AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17582 2025-11-13 cs.HC cs.GR cs.MM 80%

Crafting Dynamic Virtual Activities with Advanced Multimodal Models

Changyang Li, Qingan Yan, Minyoung Kim, Zhan Li, Yi Xu, Lap-Fai Yu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.MM

Journal ref C. Li, Q. Yan, M. Kim, Z. Li, Y. Xu and L. -F. Yu, "Crafting Dynamic Virtual Activities with Advanced Multimodal Models," 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 120-130

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17417 2025-11-13 cs.CV 70%

Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment

Robert Wijaya, Ngoc-Bao Nguyen, Ngai-Man Cheung

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21348 2025-11-13 cs.CL cs.AI 62%

Large Language Model Benchmarks in Medical Tasks

Lawrence K. Q. Yan, Qian Niu, Ming Li, Yichao Zhang, Caitlyn Heqi Yin, Cheng Fei, Benji Peng, Ziqian Bi, Pohsun Feng, Keyu Chen, Tianyang Wang, Yunze Wang, Silin Chen, Ming Liu, Junyu Liu, Xinyuan Song, Riyang Bao, Zekun Jiang, Ziyuan Qin

机构 * Hong Kong University of Science and Technology(香港科技大学) Kyoto University(京都大学) Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Cornell University(康奈尔大学) Indiana University(印第安纳大学) National Taiwan Normal University(台湾师范大学) University of Liverpool(利物浦大学) University of Edinburgh(爱丁堡大学) Zhejiang University(浙江大学) Purdue University(普渡大学) Emory University(埃默里大学) West China Biomedical Big Data Center, West China Hospital, Sichuan University(西京生物大数据中心,西京医院,四川大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 25 pages, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09064 2025-11-13 cs.CV 57%

Diversifying Counterattacks: Orthogonal Exploration for Robust CLIP Inference

Chengze Jiang, Minjing Dong, Xinli Shi, Jie Gui

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to AAAI-2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08909 2025-11-13 cs.CV 57%

Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images

Zimao Lu, Hui Xu, Bing Liu, Ke Wang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2511.09552 2025-11-13 cs.CR cs.MM 88%

Intelligent Carrier Allocation: A Cross-Modal Reasoning Framework for Adaptive Multimodal Steganography

Abhirup Das, Pranav Dudani, Shruti Sharma, Ravi Kumar C.

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.MM

Comments 8 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08642 2025-11-13 eess.IV cs.MM cs.SD 83%

Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations

Jingwen Fu, Ming Xiao, Zhonghao Lyu, Mikael Skoglund, Celimuge Wu

机构 * School of Electrical Engineering and Computer Science (EECS), KTH Royal Institute of Technology(电气工程与计算机科学学院(EECS),皇家理工学院) Department of Computer and Network Engineering, The University of Electro-Communications(计算机与网络工程系,东京电讯大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09448 2025-11-13 cs.MM cs.LG 79%

MCAD: Multimodal Context-Aware Audio Description Generation For Soccer

Lipisha Chaudhary, Trisha Mittal, Subhadra Gopalakrishnan, Ifeoma Nwogu, Jaclyn Pytlarz

机构 * University at Buffalo, SUNY(布法罗大学,SUNY) Dolby Laboratories Inc.(杜比实验室公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22229 2025-11-13 cs.SD eess.AS 79%

Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device

Zixuan Li, Xueliang Zhang, Lei Miao, Zhipeng Yan, Ying Sun, Chong Zhu

机构 * College of Computer Science, Inner Mongolia University, China(内蒙古大学计算机科学学院) Lenovo, China(联想公司)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09039 2025-11-13 cs.LG cs.CY cs.HC 78%

Fairness-Aware Few-Shot Learning for Audio-Visual Stress Detection

Anushka Sanjay Shelke, Aditya Sneh, Arya Adyasha, Haroon R. Lone

专题命中 音频语音多模态 :audio-visual(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04914 2025-11-13 cs.SD cs.AI 57%

MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages

Hardik B. Sailor, Aw Ai Ti, Chen Fang Yih Nancy, Chiu Ying Lay, Ding Yang, He Yingxu, Jiang Ridong, Li Jingtao, Liao Jingyi, Liu Zhuohan, Lu Yanfeng, Ma Yi, Manas Gupta, Muhammad Huzaifah Bin Md Shahrin, Nabilah Binte Md Johan, Nattadaporn Lertcheva, Pan Chunlei, Pham Minh Duc, Siti Maryam Binte Ahmad Subaidi, Siti Umairah Binte Mohammad Salleh, Sun Shuo, Tarun Kumar Vangani, Wang Qiongqiong, Won Cheng Yi Lewis, Wong Heng Meng Jeremy, Wu Jinyang, Zhang Huayun, Zhang Longyin, Zou Xunlong

机构 * MERaLiON Team Institute for Infocomm Research (I 2 R), A*STAR, Singapore(MERaLiON团队信息与通信研究所(I 2 R),A*STAR,新加坡)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments https://huggingface.co/MERaLiON/MERaLiON-SER-v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16774 2025-11-13 cs.CL 57%

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

Yiming Gao, Bin Wang, Chengwei Wei, Shuo Sun, AiTi Aw

机构 * Nanyang Technological University (NTU)(南洋理工大学) MiroMind(米罗Mind) Institute for Infocomm Research (I 2 R)(信息与通信研究院) A*STAR(科技研究局)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Link: https://github.com/AudioLLMs/AudioBench/tree/main/IFEval-Audio

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2505.23990 2025-11-13 cs.AI 79%

Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding

Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri, Nicholas R. Waytowich, Xiaomin Lin, Tinoosh Mohsenin

机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09342 2025-11-13 eess.SP 78%

A cross-modal pre-training framework with video data for improving performance and generalization of distributed acoustic sensing

Junyi Duan, Jiageng Chen, Zuyuan He

专题命中 视频多模态 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09137 2025-11-13 eess.SP 78%

xHAP: Cross-Modal Attention for Haptic Feedback Estimation in the Tactile Internet

Georgios Kokkinis, Alexandros Iosifidis, Qi Zhang

专题命中 视频多模态 :cross-modal(title,abstract)

Comments 12 pages, 13 figures, 3 tables, 2 algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03184 2025-11-13 eess.IV cs.CV 57%

EvRWKV: A Continuous Interactive RWKV Framework for Effective Event-Guided Low-Light Image Enhancement

Wenjie Cai, Qingguo Meng, Zhenyu Wang, Xingbo Dong, Zhe Jin

机构 * Anhui Provincial International Joint Research center for Advanced technology in Medical imaging(安徽省国际联合先进医学影像技术研究中心) School of Artificial Intelligence(人工智能学院) Anhui University(安徽大学) State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) Anhui Provincial Key Laboratory of Secure Artificial Intelligence(安徽省安全人工智能重点实验室) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(粤港澳大湾区人工智能与数字经济实验室(深圳))

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08892 2025-11-13 cs.HC cs.RO 50%

Help or Hindrance: Understanding the Impact of Robot Communication in Action Teams

Tauhid Tanjim, Jonathan St. George, Kevin Ching, Angelique Taylor

机构 * Department of Information Science at Cornell University(康奈尔大学信息科学系) Weill Cornell Medicine, Cornell University(韦尔医学院,康奈尔大学)

专题命中 视频多模态 :multimodal(abstract)

Comments This is the author's original submitted version of the paper accepted to the 2025 IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). \c{opyright} 2025 IEEE. Personal use of this material is permitted. For any other use, please contact IEEE

Journal ref 2025 34th IEEE International Conference on Robot and Human Interactive Communication RO-MAN pp. 1460-1465

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2511.09250 2025-11-13 cs.IR 82%

NeuroCLIP: Brain-Inspired Prompt Tuning for EEG-to-Image Multimodal Contrastive Learning

Jiyuan Wang, Li Zhang, Haipeng Lin, Qile Liu, Gan Huang, Ziyu Li, Zhen Liang, Xia Wu

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09228 2025-11-13 cs.CV cs.CL 62%

Taming Object Hallucinations with Verified Atomic Confidence Estimation

Jiarui Liu, Weihao Xuan, Zhijing Jin, Mona Diab

机构 * CMU(卡内基梅隆大学) The University of Tokyo(东京大学) The University of Toronto(多伦多大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 1 篇

2511.09058 2025-11-13 cs.CV 79%

VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering

Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le

机构 * Faculty Of Information Technology, VNU University of Engineering and Technology(信息科技学院,越南工程与技术大学) IT-BT Convergence Technology Division, Vietnam-Korea Institute of Science and Technology(IT-BT融合技术部,越南-韩国科学技术院) TADI Global Lab, TADI Global Company Limited(TADI全球实验室,TADI全球公司) Faculty of Finance, Banking Academy of Vietnam(金融学院,越南银行学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 3 figures, 3 tables, FAIR 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 11 篇

2506.00942 2025-11-13 cs.CL cs.AI eess.SP 84%

anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding

Haitao Li, Ziyu Li, Yiheng Mao, Ziyi Liu, Zhoujian Sun, Zhengxing Huang

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10975 2025-11-13 cs.IR cs.CL 83%

ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER

Jielong Tang, Shuang Wang, Zhenxing Wang, Jianxing Yu, Jian Yin

机构 * School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) Key Laboratory of Sustainable Tourism Smart Assessment Technology, Ministry of Culture and Tourism, Sun Yat-sen University(文化旅游可持续评估技术重点实验室,中华人民共和国文化和旅游部,中山大学) Beijing Normal University(北京师范大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments CCKS 2025 Shared Task Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10522 2025-11-13 cs.LG cs.AI cs.CV eess.AS 82%

Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction

Kaizhen Tan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09067 2025-11-13 cs.CL cs.AI 81%

MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique

Gailun Zeng, Ziyang Luo, Hongzhan Lin, Yuchen Tian, Kaixin Li, Ziyang Gong, Jianxiong Guo, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) National University of Singapore(新加坡国立大学) Beijing Normal University(北京师范大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 28 pages, 14 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24792 2025-11-13 cs.CV cs.AI 81%

PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models

Patrick Haller, Fabio Barth, Jonas Golde, Georg Rehm, Alan Akbik

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 11 tables and figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09339 2025-11-13 cs.CL 79%

mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models

Arka Mukherjee, Shreya Ghosh

机构 * Kalinga Institute of Industrial Technology (KIIT)(喀里亚理工学院) Indian Institute of Technology (IIT), Bhubaneswar(印度理工学院(班加罗尔))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to IJCNLP-AACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08901 2025-11-13 cs.CV 79%

Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency

Riling Wei, Kelu Yao, Chuanguang Yang, Jin Wang, Zhuoyan Gao, Chao Li

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07413 2025-11-13 cs.AI cs.CL cs.HC cs.LG 62%

DigiData: Training and Evaluating General-Purpose Mobile Control Agents

Yuxuan Sun, Manchen Wang, Shengyi Qian, William R. Wong, Eric Gan, Pierluca D'Oro, Alejandro Castillejo Munoz, Sneha Silwal, Pedro Matias, Nitin Kamra, Satwik Kottur, Nick Raines, Xuanyi Zhao, Joy Chen, Joseph Greer, Andrea Madotto, Allen Bolourchi, James Valori, Kevin Carlberg, Karl Ridgeway, Joseph Tighe

机构 * FAIR at Meta(Meta 的 FAIR 研究组) Meta Reality Labs(Meta 现实实验室) University of Southern California(南加州大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Website: https://facebookresearch.github.io/DigiData

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09438 2025-11-13 cs.LG cs.AI 57%

LLM-Guided Dynamic-UMAP for Personalized Federated Graph Learning

Sai Puppala, Ismail Hossain, Md Jahangir Alam, Tanzim Ahad, Sajedul Talukder

机构 * University of Texas at El Paso(德克萨斯理工大学) Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏