arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46294 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4676 篇

2511.18437 2025-11-25 cs.CV 79%

Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning

基于感知证据的强化学习用于多模态推理

Chi Zhang, Haibo Qiu, Qiming Zhang, Yufei Xu, Zhixiong Zeng, Siqi Yang, Peng Shi, Lin Ma, Jing Zhang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Meituan Inc(美团公司) The University of Sydney(悉尼大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 PEARL通过将推理锚定在验证的视觉证据上,提升多模态推理能力,有效解决视觉幻觉和奖励黑客问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18314 2025-11-25 cs.LG cs.AI 79%

AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert

AnyExperts: 多模态语言模型中基于需求的专家分配方法

Yuting Gao, Wang Lan, Hengyuan Zhao, Linjiang Huang, Si Liu, Qingpei Guo

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 AnyExperts通过按需分配专家资源,提高多模态MoE模型的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17448 2025-11-24 cs.CV 79%

MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models

MMT-ARD: 多模态多教师对抗蒸馏用于鲁棒视觉-语言模型

Yuqi Li, Junhao Dong, Chuanguang Yang, Shiping Wen, Piotr Koniusz, Tingwen Huang, Yingli Tian, Yew-Soon Ong

机构 * The City University of New York, CUNY(纽约城市大学) Nanyang Technological University(南洋理工大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Technology Sydney(悉尼技术大学) Data61, CSIRO(CSIRO数据61研究所) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 MMT-ARD通过多教师对抗蒸馏提升视觉-语言模型的对抗鲁棒性,实验显示鲁棒精度提升4.32%,训练效率提高2.3倍。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11313 2025-11-24 cs.CV 79%

DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding

DocSLM:一种用于长多模态文档理解的小型视觉-语言模型

Tanveer Hannan, Dimitrios Mallios, Parth Pathak, Faegheh Sardari, Thomas Seidl, Gedas Bertasius, Mohsen Fayyaz, Sunando Sengupta

机构 * Microsoft(微软公司) LMU Munich(慕尼黑大学) MCML FAIR Meta(Meta FAIR) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 DocSLM通过高效的小型视觉-语言模型,在受限内存下实现长多模态文档理解,使用更少的资源达到与先进方法相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07864 2025-11-18 cs.CV 79%

Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization

Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu

机构 * MBZUAI Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院) University of Copenhagen(哥本哈根大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10974 2025-11-17 cs.CV 79%

Preserving Cross-Modal Consistency for CLIP-based Class-Incremental Learning

Haoran Chen, Houze Xu, Micah Goldblum, Daoguo Dong, Zuxuan Wu

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Columbia University(哥伦比亚大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10074 2025-11-14 cs.CV cs.SY eess.SY 79%

VLF-MSC: Vision-Language Feature-Based Multimodal Semantic Communication System

Gwangyeon Ahn, Jiwan Seo, Joonhyuk Kang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments To appear in the AI4NextG Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19110 2025-11-14 cs.CV 79%

LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models

Zhihui Guo, Xin Man, Hui Xu, Jie Shao, Zhiguo Jiang, Xianchao Zhang, Heng Tao Shen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07966 2025-11-12 cs.CV 79%

Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection

Shenao Zhao, Pengpeng Liang, Zhoufan Yang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09997 2025-11-11 cs.CV 79%

Descriptive Image-Text Matching with Graded Contextual Similarity

Jinhyun Jang, Jiyoung Lee, Kwanghoon Sohn

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments This version is incomplete and requires substantial revisions and extensions. We withdraw the paper and plan to submit a thoroughly revised version as a new submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20013 2025-11-11 cs.LG cs.AI cs.IR 79%

Cross-Platform E-Commerce Product Categorization and Recategorization: A Multimodal Hierarchical Classification Approach

Lotte Gross, Rebecca Walter, Nicole Zoppi, Adrien Justus, Alessandro Gambetti, Qiwei Han, Maximilian Kaiser

机构 * Nova School of Business and Economics(诺瓦商业与经济学院) Nova School of Business(诺瓦商业学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accetped at IEEE BigData 2025, 10 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23769 2025-11-07 cs.CV 79%

TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu, Zhen Zhu, Derek Hoiem

机构 * Siebel School of Computing and Data Science(塞比尔计算与数据科学学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Published in TMLR, with a J2C Certification

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00916 2025-11-04 cs.CV 79%

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

Yan Shu, Chi Liu, Robin Chen, Derek Li, Bryan Dai

机构 * Ubiquant(乌比量化)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00821 2025-11-04 cs.CV 79%

OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models

Ruoxiang Huang, Xindian Ma, Rundong Kong, Zhen Yuan, Peng Zhang

机构 * Tianjin University(天津大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02095 2025-11-04 cs.CV cs.LG 79%

Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences

Hyojin Bahng, Caroline Chan, Fredo Durand, Phillip Isola

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01341 2025-11-04 cs.CL 79%

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow York University(约克大学) Mila – Quebec AI Institute(魁北克人工智能研究院) École de Technologie Supérieure(魁北克高等技术学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) CIFAR AI Chair(CIFAR人工智能 chair) Polytechnique Montréal(蒙特利尔理工学院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24331 2025-10-29 cs.LG cs.CV 79%

What do vision-language models see in the context? Investigating multimodal in-context learning

Gabriel O. dos Santos, Esther Colombini, Sandra Avila

机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21093 2025-10-27 cs.AI 79%

MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning

Siyong Chen, Jinbo Wen, Jiawen Kang, Tenghui Huang, Xumin Huang, Yuanjia Su, Hudan Pan, Zishao Zhong, Dusit Niyato, Shengli Xie, Dong In Kim

机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院) College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) State Key Laboratory of Traditional Chinese Medicine Syndrome, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangdong Provincial Hospital of Chinese Medicine, Guangdong Provincial Academy of Chinese Medical Sciences(广东省中医药科学院中医证候重点实验室,广州中医药大学第二附属医院,广东省中医院,广东省中医药科学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Department of Electrical and Computer Engineering, Sungkyunkwan University(成均馆大学电子与计算机工程系)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16198 2025-10-21 cs.CL 79%

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad, Hesham Abdelgawad, Mohamed Aref, Ahmed Fares

机构 * Department of Electrical Engineering, Faculty of Engineering at Shoubra, Benha University, Cairo 11629, Egypt(电气工程系,谢布拉工程学院,本海大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16036 2025-10-21 cs.CV 79%

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

Zewen Li, Zitong Yu, Qilang Ye, Weicheng Xie, Wei Zhuo, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院) College of Computer Science, Nankai University(南开大学计算机学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室) National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Instrumentation and Measurement (TIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12974 2025-10-20 cs.CV 79%

Scope: Selective Cross-modal Orchestration of Visual Perception Experts

Tianyu Zhang, Suyuchen Wang, Chao Wang, Juan Rodriguez, Ahmed Masry, Xiangru Jian, Yoshua Bengio, Perouz Taslakian

机构 * ServiceNow Université de Montréal(蒙特利尔大学) École de Technologie Supérieure(高级技术学院) University of Waterloo(滑铁卢大学) McGill University(麦吉尔大学) York University(约克大学) CIFAR AI Chair(CIFAR人工智能主席) Mila Law Zero

专题命中 图文多模态 :cross-modal(title);image-text(abstract);分类 cs.CV

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25373 2025-10-17 cs.AI 79%

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

Chenyue Zhou, Mingxuan Wang, Yanbiao Ma, Chenxu Wu, Wanyi Chen, Zhe Qian, Xinyu Liu, Yiwei Zhang, Junhao Wang, Hengbo Xu, Fei Luo, Xiaohua Chen, Xiaoshuai Hao, Hehan Li, Andi Zhang, Wenxuan Wang, Kaiyan Zhang, Guoli Jia, Lingling Li, Zhiwu Lu, Yang Lu, Yike Guo

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Xiamen University(厦门大学) The Hong Kong University of Science and Technology(香港理工大学) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10342 2025-10-14 cs.CV 79%

Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis

Yu-Hsuan Lin

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 7 pages, 4 figures. Preprint submitted to arXiv in October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10117 2025-10-14 cs.AI 79%

DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay

Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu, Baixuan Xu, Yauwai Yim, Chunkit Chan, Jiaxin Bai, Yangqiu Song

机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(计算机科学与工程系,香港科技大学,香港特别行政区,中国)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments EMNLP 2025 Wordplay (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10104 2025-10-14 cs.CV 79%

Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models

Minbin Huang, Runhui Huang, Chuanyang Zheng, Jingyao Li, Guoxuan Chen, Han Shi, Hong Cheng

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09358 2025-10-13 cs.CV 79%

Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models

Qihang Ma, Shengyu Li, Jie Tang, Dingkang Yang, Shaodong Chen, Yingyi Zhang, Chao Feng, Jiao Ran

机构 * ByteDance Douyin Content Group(字节跳动抖音内容团队)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments EMNLP2025. Code is avaible at https://github.com/bytedance/DynamicCoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06092 2025-10-10 cs.CV 79%

Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation

Yachun Mi, Yu Li, Yanting Li, Chen Hui, Tong Zhang, Zhixuan Li, Chenyue Song, Wei Yang Bryan Lim, Shaohui Liu

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18743 2025-10-07 cs.CV 79%

SAR-TEXT: A Large-Scale SAR Image-Text Dataset Built with SAR-Narrator and A Progressive Learning Strategy for Downstream Tasks

Yiguo He, Xinjun Cheng, Junjie Zhu, Chunping Qiu, Jun Wang, Xichuan Zhang, Qiangjuan Huang, Ke Yang

机构 * Intelligent Game and Decision Lab(智能游戏与决策实验室)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments IEEE Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25791 2025-10-01 cs.CV 79%

EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks

Yuan Gao, Sangwook Kim, Chris McIntosh

机构 * Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络) Department of Medical Biophysics, UofT(医学生物物理学系) Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络) Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学) Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute) Department of Medical Imaging, UofT(医学影像学系) Vector Institute, Toronto(向量研究所)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments MICCAI 2025

Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. MICCAI 2025. Lecture Notes in Computer Science, vol 15964. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18839 2025-09-24 cs.CV 79%

Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography

Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza

机构 * Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏