arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-16 至 2025-10-16 共收录 52 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 10 篇

2507.10013 2025-10-16 cs.CV cs.CL 84%

Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect

Tom Kouwenhoven, Kiana Shahrasbi, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science(莱顿先进计算机科学研究所) Leiden University(莱顿大学)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13497 2025-10-16 cs.LG cs.AI 83%

DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation

Zexin Wang, Lin Shi, Haoyu Wu, Junru Luo, Xiangzeng Kong, Jun Qi

机构 * Aliyun School of Big Data, Changzhou University(阿里云大数据学院,长洲大学) Department of Computing, Xi’an JiaoTong-Liverpool University(计算系,西安交通大学-利物浦大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Center for Artificial Intelligence in Agriculture, Fujian Agriculture and Forestry University(农业人工智能中心,福建农林大学)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14138 2025-10-16 cs.CV cs.AI 81%

ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom

Jingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu, Lei Li, Jiahui Gao, Jiyue Jiang, Lingpeng Kong, Chuan Wu

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13364 2025-10-16 cs.CV cs.AI 76%

Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity

MingZe Tang, Jubal Chandy Jacob

机构 * Department of Computing Science University of Aberdeen(计算科学系阿伯丁大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12845 2025-10-16 cs.CL cs.AI cs.CV cs.RO 75%

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

Jesse Atuhurra, Iqra Ali, Tomoya Iwakura, Hidetaka Kamigaito, Tatsuya Hiraoka

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13235 2025-10-16 cs.CV 70%

EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking

Yukuan Zhang, Jiarui Zhao, Shangqing Nie, Jin Kuang, Shengsheng Wang

机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(吉林大学教育部长春符号计算与知识工程重点实验室) Yangtze University(扬子大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12931 2025-10-16 cs.CV cs.CL 62%

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

Sanghyun Byun, Jung Ick Guack, Mohanad Odema, Baisub Lee, Jacob Song, Woo Seong Chung

机构 * LG Electronics USA(LG电子美国公司)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted to PMLR and NeurIPS 2025 UniReps

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15298 2025-10-16 cs.CV cs.MM 62%

MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering

Xinqi Fan, Jingting Li, John See, Moi Hoon Yap, Wen-Huang Cheng, Xiaobai Li, Xiaopeng Hong, Su-Jing Wang, Adrian K. Davision

机构 * Department of Computing and Mathematics, Manchester Metropolitan University(计算与数学系,曼彻斯特 Metropolitan 大学) State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences(认知科学与心理健康国家重点实验室,心理学研究所,中国科学院) Department of Psychology, University of the Chinese Academy of Sciences(心理学系,中国科学院大学) National Taiwan University(台湾大学) Zhejiang University(浙江大学) University of Oulu(奥卢大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Micro-Expression Grand Challenge (MEGC) at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13359 2025-10-16 cs.IR cs.CV cs.LG 57%

Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models

Yuki Yada, Sho Akiyama, Ryo Watanabe, Yuta Ueno, Yusuke Shido, Andre Rusli

机构 * Mercari, Inc.(Mercari公司)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to ACM RecSys 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13190 2025-10-16 cs.CL 57%

SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 9 篇

2507.09945 2025-10-16 cs.MM cs.CV 86%

ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization

Huilai Li, Yonghao Dang, Ying Xing, Yiming Wang, Jianqin Yin

机构 * School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(智能工程与自动化学院,北京邮电大学) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13182 2025-10-16 cs.LG 82%

Information-Theoretic Criteria for Knowledge Distillation in Multimodal Learning

Rongrong Xie, Yizhou Xu, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati (SISSA)(国际先进研究高等学院) École Polytechnique Fédérale de Lausanne (EPFL)(日内瓦联邦理工学院)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13281 2025-10-16 eess.AS cs.CL cs.LG 81%

Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses

Sungnyun Kim, Kangwook Jang, Sungwoo Cho, Joon Son Chung, Hoirin Kim, Se-Young Yun

机构 * Kim Jaechul Graduate School of AI, KAIST(金 Jaechul人工智能研究生院,韩国科学技术院) School of Electrical Engineering, KAIST(电气工程学院,韩国科学技术院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL、eess.AS

Comments Preprint work

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13308 2025-10-16 eess.AS 79%

Towards Multimodal Query-Based Spatial Audio Source Extraction

Chenxin Yu, Hao Ma, Xu Li, Xiao-Lei Zhang, Mingjie Shao, Chi Zhang, Xuelong Li

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12851 2025-10-16 cs.SD cs.LG eess.AS 79%

Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models

Tsung-En Lin, Kuan-Yi Lee, Hung-Yi Lee

机构 * National Taiwan University(国立台湾大学) ASUS Open Cloud Infrastructure Software Center(ASUS开放云基础设施软件中心)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 eess.AS

Comments Note: This preprint is a version of the paper submitted to ICASSP 2026. The author list here includes contributors who provided additional supervision and guidance. The official ICASSP submission may differ slightly in author composition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13351 2025-10-16 cs.CL cs.AI 62%

Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems

Karthik Avinash, Nikhil Pareek, Rishav Hada

机构 * FutureAGI Inc.(未来人工智能公司)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13344 2025-10-16 cs.SD cs.CL 57%

UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE

Zhenyu Liu, Yunxin Li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Jinchao Li, Qi Wang, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Min Zhang

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05803 2025-10-16 cs.GR cs.CV 57%

PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis

Yihuan Huang, Jiajun Liu, Yanzhen Ren, Jun Xue, Wuyang Liu, Zongkun Sun

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航空航天信息安全部分和可信计算重点实验室、教育部、网络安全科学与工程学院、武汉大学) School of Cyber Science and Engineering, Wuhan University(网络安全科学与工程学院、武汉大学) School of Police Information, Shandong Police College(警务信息学院、山东警察学院)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13558 2025-10-16 cs.SD 50%

Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module

Ruitao Feng, Bixi Zhang, Sheng Liang, Zheng Yuan

机构 * The University of Hong Kong, Fauclty of Science, Hong Kong(香港大学科学学院) Aix-Marseille University, Laboratoire Parole et Langage (LPL), France(艾克斯-马赛大学语言与言语实验室(LPL))

专题命中 音频语音多模态 :multimodal(abstract)

Comments 5 pages, 1 figures. Code is available at: https://github.com/forfrt/SteerMoE. Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 3 篇

2507.05258 2025-10-16 cs.CV cs.LG 70%

Spatio-Temporal LLM: Reasoning about Environments and Actions

Haozhen Zheng, Beitong Tian, Mingyuan Wu, Zhenggang Tang, Klara Nahrstedt, Alex Schwing

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Code and data are available at https://zoezheng126.github.io/STLLM-website/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16836 2025-10-16 cs.CV cs.AI 62%

Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning

Fanrui Zhang, Dian Li, Qiang Zhang, Jun Chen, Gang Liu, Junxiong Lin, Jiahong Yan, Jiawei Liu, Zheng-Jun Zha

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, USTC(脑启发智能感知与认知国家重点实验室,中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Tencent QQ(腾讯QQ) Fudan University(复旦大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 34 pages, 25 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13019 2025-10-16 physics.optics 50%

Phase Matching of Orbital Angular Momentum in Rare Earth Ion Doped Solid State Systems

Owen R. Wolfe, Joshua Dugre, Grant Kirkland, R. Krishna Mohan

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2510.13366 2025-10-16 cs.CL cs.AI 62%

Document Intelligence in the Era of Large Language Models: A Survey

Weishi Wang, Hengchang Hu, Zhijie Zhang, Zhaochen Li, Hongxin Shao, Daniel Dahlmeier

机构 * SAP, Singapore(新加坡SAP)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13406 2025-10-16 cs.LG 50%

When Embedding Models Meet: Procrustes Bounds and Applications

Lucas Maystre, Alvaro Ortega Gonzalez, Charles Park, Rares Dolga, Tudor Berariu, Yu Zhao, Kamil Ciosek

机构 * UiPath Spotify

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 2 篇

2510.13804 2025-10-16 cs.CV cs.AI cs.CL 82%

Generative Universal Verifier as Multimodal Meta-Reasoner

Xinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao, Renrui Zhang, Ruihang Chu, Ling Yang, Yujiu Yang

机构 * Tsinghua University(清华大学) ByteDance Seed(字节跳动种子) Princeton University(普林斯顿大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12041 2025-10-16 cs.CL 57%

Improving Text-to-Image Generation with Input-Side Inference-Time Scaling

Ruibo Chen, Jiacheng Pan, Heng Huang, Zhenheng Yang

机构 * TikTok University of Maryland, College Park(马里兰大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 13 篇

2402.03173 2025-10-16 cs.CL cs.AI cs.CV 87%

MULTI: Multimodal Understanding Leaderboard with Text and Images

Zichen Zhu, Yang Xu, Lu Chen, Jingkai Yang, Yichuan Ma, Yiming Sun, Hailin Wen, Jiaqi Liu, Jinyu Cai, Yingzi Ma, Situo Zhang, Zihan Zhao, Liangtai Sun, Kai Yu

机构 * X-LANCE Lab, School of Computer Science, Key Laboratory of Artificial Intelligence\ of Education, Shanghai Jiao Tong University, Shanghai 200240 , China Jiangsu Key Lab of Language Computing, Suzhou 215123 , China College of Computing Data Science, Nanyang Technological University, Singapore 639798 , Singapore Suzhou Laboratory, Suzhou 215123 , China

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 24 pages, 19 figures, 10 tables. Details and access are available at: https://OpenDFM.github.io/MULTI-Benchmark/

Journal ref Sci. China Inf. Sci. 68, 200107 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19330 2025-10-16 eess.SP cs.AI cs.HC cs.LG cs.MM 81%

LibEMER: A novel benchmark and algorithms library for EEG-based Multimodal Emotion Recognition

Zejun Liu, Yunshan Chen, Chengxi Xie, Yugui Xie, Huan Liu

机构 * XJTU-POLIMI Joint School, Xi'an Jiaotong University School of Computer Science Technology, Xi'an Jiaotong University MIGU Video Co., Ltd., Shanghai, China

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10412 2025-10-16 cs.LG cs.AI cs.CL 81%

Time-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series

Ching Chang, Jeehyun Hwang, Yidan Shi, Haixin Wang, Wen-Chih Peng, Tien-Fu Chen, Wei Wang

机构 * University of California, Los Angeles(加州大学洛杉矶分校) National Yang Ming Chiao Tung University(国立阳明交通大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted by the NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13211 2025-10-16 cs.CV cs.AI 81%

MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance

Subin Kim, Hoonrae Kim, Jihyun Lee, Yejin Jeon, Gary Geunbae Lee

机构 * KT Corporation, Republic of Korea(韩国KT公司) Graduate School of Artificial Intelligence, POSTECH, Republic of Korea(POSTECH人工智能研究生院) Computer Science and Engineering, POSTECH, Republic of Korea(POSTECH计算机科学与工程系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏