arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-20 至 2025-10-20 共收录 52 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2505.14714 2025-10-20 cs.CV cs.AI cs.CL 87%

KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection

Tuan-Vinh La, Minh-Hieu Nguyen, Minh-Son Dao

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Withdrawn by the authors due to lack of explicit agreement from all co-authors to post this version publicly on arXiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15162 2025-10-20 cs.CV cs.CL 86%

Train a Unified Multimodal Data Quality Classifier with Synthetic Data

Weizhi Wang, Rongmei Lin, Shiyang Li, Colin Lockard, Ritesh Sarkhel, Sanket Lokegaonkar, Jingbo Shang, Xifeng Yan, Nasser Zalmout, Xian Li

机构 * UC Santa Barbara(加州大学圣芭芭拉分校) Amazon Stores Foundational AI(亚马逊商店基础人工智能) UC San Diego(加州大学圣地亚哥分校)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12974 2025-10-20 cs.CV 79%

Scope: Selective Cross-modal Orchestration of Visual Perception Experts

Tianyu Zhang, Suyuchen Wang, Chao Wang, Juan Rodriguez, Ahmed Masry, Xiangru Jian, Yoshua Bengio, Perouz Taslakian

机构 * ServiceNow Université de Montréal(蒙特利尔大学) École de Technologie Supérieure(高级技术学院) University of Waterloo(滑铁卢大学) McGill University(麦吉尔大学) York University(约克大学) CIFAR AI Chair(CIFAR人工智能主席) Mila Law Zero

专题命中 图文多模态 :cross-modal(title);image-text(abstract);分类 cs.CV

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05342 2025-10-20 cs.CV cs.AI 62%

Refer to Any Segmentation Mask Group With Vision-Language Prompts

Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu, Lingzhi Zhang, Jiuxiang Gu, HyunJoon Jung, Liang-Yan Gui, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Adobe

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17092 2025-10-20 cs.CV cs.CL 62%

Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI

Syed Abdul Gaffar Shakhadri, Kruthika KR, Kartik Basavaraj Angadi

机构 * SandLogic Technologies Pvt Ltd(沙德逻辑技术 Pvt Ltd)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09607 2025-10-20 cs.CV 57%

VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation

Shaoqi Dong, Chaoyou Fu, Haihan Gao, Yi-Fan Zhang, Chi Yan, Chu Wu, Xiaoyu Liu, Yunhang Shen, Jing Huo, Deqiang Jiang, Haoyu Cao, Yang Gao, Xing Sun, Ran He, Caifeng Shan

机构 * Nanjing University(南京大学) Tencent Youtu Lab(腾讯优图实验室) CASIA(中国科学院自动化研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Homepage: https://ltbai.github.io/VITA-VLA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01842 2025-10-20 cs.LG cs.AI 57%

GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping

Qifu Wen, Xi Zeng, Zihan Zhou, Shuaijun Liu, Mehdi Hosseinzadeh, Ningxin Su, Reza Rawassizadeh

机构 * Department of Computer Science, Boston University Metropolitan College(波士顿大学计算机科学系) Information Hub, The Hong Kong University of Science and Technology, Guangzhou(香港科技大学广州信息中心) School of Engineering and Technology, Duy Tan University, Da Nang, Vietnam(杜益大学工程科技学院,岘港,越南) Department of AI, School of Computer Science and Engineering, Galgotias University, Greater Noida, India(加洛吉亚大学人工智能系,诺伊达,印度)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15508 2025-10-20 cs.LG cs.NA math.NA stat.ML 50%

Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity

Naoki Yoshida, Satoshi Hayakawa, Yuhta Takida, Toshimitsu Uesaka, Hiromi Wakaki, Yuki Mitsufuji

机构 * The University of Tokyo(东京大学) Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能)

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2510.15685 2025-10-20 cs.CL 79%

Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection

Joshua Wolfe Brook, Ilia Markov

机构 * Computational Linguistics \& Text Mining Lab Vrije Universiteit Amsterdam

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 8 pages, 9 figures, submitted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10396 2025-10-20 cs.SD 78%

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations

Wenxiang Guo, Changhao Pan, Zhiyuan Zhu, Xintong Hu, Yu Zhang, Li Tang, Rui Yang, Han Wang, Zongbao Zhang, Yuhan Wang, Yixuan Chen, Hankun Xu, Ke Xu, Pengfei Fan, Zhetao Chen, Yanhao Yu, Qiange Huang, Fei Wu, Zhou Zhao

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10444 2025-10-20 cs.CL cs.AI 62%

Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance

Jingyi Chen, Zhimeng Guo, Jiyun Chun, Pichao Wang, Andrew Perrault, Micha Elsner

机构 * Department of Linguistics, The Ohio State University, USA(语言学系,俄亥俄州立大学) Department of Computer Science and Engineering, The Ohio State University, USA(计算机科学与工程系,俄亥俄州立大学) Department of Information Sciences and Technology, Penn State University, USA(信息科学与技术系,宾夕法尼亚州立大学) Amazon, USA(亚马逊公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15757 2025-10-20 cs.LG cs.CV cs.NE 57%

Poultry Farm Intelligence: An Integrated Multi-Sensor AI Platform for Enhanced Welfare and Productivity

Pieris Panagi, Savvas Karatsiolis, Kyriacos Mosphilis, Nicholas Hadjisavvas, Andreas Kamilaris, Nicolas Nicolaou, Efstathios Stavrakis, Vassilis Vassiliades

机构 * CYENS - Centre of Excellence(CYENS 卓越中心) Algolysis Ltd(Algolysis 公司) Department of Computer Science(计算机科学系)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10248 2025-10-20 cs.LG cs.AI 57%

Reasoning-Enhanced Large Language Models for Molecular Property Prediction

Jiaxi Zhuang, Yaorui Shi, Jue Hou, Yunong He, Mingwei Ye, Mingjun Xu, Yuming Su, Linfeng Zhang, Ying Qian, Linfeng Zhang, Guolin Ke, Hengxing Cai

机构 * DP Technology(DP技术公司) Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02569 2025-10-20 cs.CL 57%

Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models

Tolúlopé Ògúnrèmí, Christopher D. Manning, Dan Jurafsky, Karen Livescu

机构 * Stanford University(斯坦福大学) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00059 2025-10-20 cs.CV cs.LG 57%

Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models

Rui Hu, Delai Qiu, Shuyu Wei, Jiaming Zhang, Yining Wang, Shengping Liu, Jitao Sang

机构 * Beijing Key Lab of Traffic Data Analysis and Mining(北京交通数据挖掘重点实验室) Beijing Jiaotong University(北京交通大学) Unisound AI Technology Co., Ltd.(Unisound人工智能技术有限公司) Peng Cheng Lab(鹏城实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACL 2025 Findings

Journal ref Findings of the Association for Computational Linguistics: ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 1 篇

2510.14992 2025-10-20 cs.CV cs.AI 73%

GAZE:Governance-Aware pre-annotation for Zero-shot World Model Environments

Leela Krishna, Mengyang Zhao, Saicharithreddy Pasula, Harshit Rajgarhia, Abhishek Mukherji

机构 * Centific Global Solutions Inc.(Centific全球解决方案公司)

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2505.19312 2025-10-20 cs.IR 85%

DocMMIR: A Framework for Document Multi-modal Information Retrieval

Zirui Li, Siwei Wu, Yizhi Li, Xingyu Wang, Yi Zhou, Chenghua Lin

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)

Comments Accepted for publication at EMNLP 2025 Findings. Code and data publicly available at https://github.com/J1mL1/DocMMIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15543 2025-10-20 cs.CL cs.AI cs.IR cs.MM 82%

MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval

Qiyu Wu, Shuyang Cui, Satoshi Hayakawa, Wei-Yao Wang, Hiromi Wakaki, Yuki Mitsufuji

机构 * Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15547 2025-10-20 cs.AI cs.ET cs.LG cs.SY eess.SP eess.SY 79%

Hypergraph Contrastive Sensor Fusion for Multimodal Fault Diagnosis in Induction Motors

Usman Ali, Ali Zia, Waqas Ali, Umer Ramzan, Abdul Rehman, Muhammad Tayyab Chaudhry, Wei Xiang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Submitted to IEEE Sensors Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09585 2025-10-20 cs.CV 79%

Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation

Jitesh Jain, Zhengyuan Yang, Humphrey Shi, Jianfeng Gao, Jianwei Yang

机构 * Microsoft Research, Redmond(微软研究院(红mond)) Meta Superintelligence Labs(Meta超智能实验室)

专题命中 跨模态检索 :multimodal(title);MLLM(abstract);分类 cs.CV

Comments Project Page: https://praeclarumjj3.github.io/visper_lm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23379 2025-10-20 cs.CL cs.AI cs.CV 75%

CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding

Xi Zhang, Zaiqiao Meng, Jake Lever, Edmond S. L. Ho

机构 * School of Computing Science, University of Glasgow(计算科学学院,格拉斯哥大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Preprint, 27 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 6 篇

2510.15068 2025-10-20 cs.CR cs.AI 83%

Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling

Deyue Zhang, Dongdong Yang, Junjie Mu, Quancheng Zou, Zonghao Ying, Wenzhuo Xu, Zhao Liu, Xuan Wang, Xiangzheng Zhang

机构 * AI Security Lab(360人工智能安全实验室) Politecnico di Milano(米兰理工大学) Beihang University(北京航空航天大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13253 2025-10-20 cs.CV cs.AI cs.LG 81%

End-to-End Multi-Modal Diffusion Mamba

Chunhao Lu, Qiang Lu, Meichen Dong, Jake Luo

机构 * China University of Petroleum-Beijing(中国石油大学(北京)) Leyard Optoelectronic(莱亚德光电) University of Wisconsin-Milwaukee(威斯康星大学密尔沃基分校)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15176 2025-10-20 cs.CV cs.AI 62%

Methods and Trends in Detecting AI-Generated Images: A Comprehensive Review

Arpan Mahara, Naphtali Rishe

机构 * Knight Foundation School of Computing and Information Sciences, Florida International University(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 34 pages, 4 Figures, 10 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15857 2025-10-20 cs.CV 57%

BLIP3o-NEXT: Next Frontier of Native Image Generation

Jiuhai Chen, Le Xue, Zhiyang Xu, Xichen Pan, Shusheng Yang, Can Qin, An Yan, Honglu Zhou, Zeyuan Chen, Lifu Huang, Tianyi Zhou, Junnan Li, Silvio Savarese, Caiming Xiong, Ran Xu

机构 * Salesforce Research(Salesforce研究部) University of Maryland(马里兰大学) Virginia Tech(弗吉尼亚理工大学) New York University(纽约大学) UC Davis(加州大学戴维斯分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03409 2025-10-20 cs.CV 57%

PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation

Ao Wang, Hui Chen, Jiaxin Li, Jianchao Tan, Kefeng Zhang, Xunliang Cai, Zijia Lin, Jungong Han, Guiguang Ding

机构 * School of Software, Tsinghua University(清华大学软件学院) BNRist, Tsinghua University(清华大学BNRist) Meituan Inc.(美团公司) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025 Camera-ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16039 2025-10-20 physics.med-ph 50%

Label-Free Intraoperative Imaging of Hemodynamics using Deep Learning

Yan Shi, Denghui Zhao, Jingyi Yu, Wei Ni, Pengcheng Li, Yun Gu, Peng Miao, Shanbao Tong

专题命中 多模态生成 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 10 篇

2510.15595 2025-10-20 cs.CV 83%

FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification

Zhen Sun, Lei Tan, Yunhang Shen, Chengmao Cai, Xing Sun, Pingyang Dai, Liujuan Cao, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室,中国教育部,厦门大学) National University of Singapore(新加坡国立大学) Tencent YouTu Lab(腾讯优图实验室) Institute of Artificial Intelligence,Xiamen University(人工智能研究院,厦门大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15684 2025-10-20 cs.CV cs.AI 81%

Towards Label-Free Brain Tumor Segmentation: Unsupervised Learning with Multimodal MRI

Gerard Comas-Quiles, Carles Garcia-Cabrera, Julia Dietlmeier, Noel E. O'Connor, Ferran Marques

机构 * Universitat Politècnica de Catalunya (UPC)(西班牙巴塞罗那理工大学) University College Dublin (UCD)(都柏林大学) Dublin City University (DCU)(都柏林城市大学) Insight Research Ireland Center For Data Analytics(爱尔兰数据分析研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5 figures, BraTS GoAT 2025 challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08636 2025-10-20 cs.CV cs.AI 81%

Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models

Xingrui Wang, Wufei Ma, Tiezheng Zhang, Celso M de Melo, Jieneng Chen, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in CVPR 2025 as Highlight. Data and code are released at https://github.com/XingruiWang/Spatial457

详情

展开后加载摘要…

URL PDF HTML 收藏