arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4672 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4672 篇

2511.11313 2025-11-24 cs.CV 79%

DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding

DocSLM:一种用于长多模态文档理解的小型视觉-语言模型

Tanveer Hannan, Dimitrios Mallios, Parth Pathak, Faegheh Sardari, Thomas Seidl, Gedas Bertasius, Mohsen Fayyaz, Sunando Sengupta

机构 * Microsoft(微软公司) LMU Munich(慕尼黑大学) MCML FAIR Meta(Meta FAIR) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 DocSLM通过高效的小型视觉-语言模型,在受限内存下实现长多模态文档理解,使用更少的资源达到与先进方法相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07864 2025-11-18 cs.CV 79%

Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization

Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu

机构 * MBZUAI Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院) University of Copenhagen(哥本哈根大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10974 2025-11-17 cs.CV 79%

Preserving Cross-Modal Consistency for CLIP-based Class-Incremental Learning

Haoran Chen, Houze Xu, Micah Goldblum, Daoguo Dong, Zuxuan Wu

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Columbia University(哥伦比亚大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10074 2025-11-14 cs.CV cs.SY eess.SY 79%

VLF-MSC: Vision-Language Feature-Based Multimodal Semantic Communication System

Gwangyeon Ahn, Jiwan Seo, Joonhyuk Kang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments To appear in the AI4NextG Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19110 2025-11-14 cs.CV 79%

LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models

Zhihui Guo, Xin Man, Hui Xu, Jie Shao, Zhiguo Jiang, Xianchao Zhang, Heng Tao Shen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07966 2025-11-12 cs.CV 79%

Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection

Shenao Zhao, Pengpeng Liang, Zhoufan Yang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09997 2025-11-11 cs.CV 79%

Descriptive Image-Text Matching with Graded Contextual Similarity

Jinhyun Jang, Jiyoung Lee, Kwanghoon Sohn

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments This version is incomplete and requires substantial revisions and extensions. We withdraw the paper and plan to submit a thoroughly revised version as a new submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20013 2025-11-11 cs.LG cs.AI cs.IR 79%

Cross-Platform E-Commerce Product Categorization and Recategorization: A Multimodal Hierarchical Classification Approach

Lotte Gross, Rebecca Walter, Nicole Zoppi, Adrien Justus, Alessandro Gambetti, Qiwei Han, Maximilian Kaiser

机构 * Nova School of Business and Economics(诺瓦商业与经济学院) Nova School of Business(诺瓦商业学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accetped at IEEE BigData 2025, 10 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23769 2025-11-07 cs.CV 79%

TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu, Zhen Zhu, Derek Hoiem

机构 * Siebel School of Computing and Data Science(塞比尔计算与数据科学学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Published in TMLR, with a J2C Certification

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00916 2025-11-04 cs.CV 79%

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

Yan Shu, Chi Liu, Robin Chen, Derek Li, Bryan Dai

机构 * Ubiquant(乌比量化)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00821 2025-11-04 cs.CV 79%

OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models

Ruoxiang Huang, Xindian Ma, Rundong Kong, Zhen Yuan, Peng Zhang

机构 * Tianjin University(天津大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02095 2025-11-04 cs.CV cs.LG 79%

Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences

Hyojin Bahng, Caroline Chan, Fredo Durand, Phillip Isola

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01341 2025-11-04 cs.CL 79%

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow York University(约克大学) Mila – Quebec AI Institute(魁北克人工智能研究院) École de Technologie Supérieure(魁北克高等技术学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) CIFAR AI Chair(CIFAR人工智能 chair) Polytechnique Montréal(蒙特利尔理工学院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24331 2025-10-29 cs.LG cs.CV 79%

What do vision-language models see in the context? Investigating multimodal in-context learning

Gabriel O. dos Santos, Esther Colombini, Sandra Avila

机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21093 2025-10-27 cs.AI 79%

MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning

Siyong Chen, Jinbo Wen, Jiawen Kang, Tenghui Huang, Xumin Huang, Yuanjia Su, Hudan Pan, Zishao Zhong, Dusit Niyato, Shengli Xie, Dong In Kim

机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院) College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) State Key Laboratory of Traditional Chinese Medicine Syndrome, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangdong Provincial Hospital of Chinese Medicine, Guangdong Provincial Academy of Chinese Medical Sciences(广东省中医药科学院中医证候重点实验室,广州中医药大学第二附属医院,广东省中医院,广东省中医药科学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Department of Electrical and Computer Engineering, Sungkyunkwan University(成均馆大学电子与计算机工程系)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16198 2025-10-21 cs.CL 79%

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad, Hesham Abdelgawad, Mohamed Aref, Ahmed Fares

机构 * Department of Electrical Engineering, Faculty of Engineering at Shoubra, Benha University, Cairo 11629, Egypt(电气工程系,谢布拉工程学院,本海大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16036 2025-10-21 cs.CV 79%

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

Zewen Li, Zitong Yu, Qilang Ye, Weicheng Xie, Wei Zhuo, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院) College of Computer Science, Nankai University(南开大学计算机学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室) National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Instrumentation and Measurement (TIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12974 2025-10-20 cs.CV 79%

Scope: Selective Cross-modal Orchestration of Visual Perception Experts

Tianyu Zhang, Suyuchen Wang, Chao Wang, Juan Rodriguez, Ahmed Masry, Xiangru Jian, Yoshua Bengio, Perouz Taslakian

机构 * ServiceNow Université de Montréal(蒙特利尔大学) École de Technologie Supérieure(高级技术学院) University of Waterloo(滑铁卢大学) McGill University(麦吉尔大学) York University(约克大学) CIFAR AI Chair(CIFAR人工智能主席) Mila Law Zero

专题命中 图文多模态 :cross-modal(title);image-text(abstract);分类 cs.CV

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25373 2025-10-17 cs.AI 79%

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

Chenyue Zhou, Mingxuan Wang, Yanbiao Ma, Chenxu Wu, Wanyi Chen, Zhe Qian, Xinyu Liu, Yiwei Zhang, Junhao Wang, Hengbo Xu, Fei Luo, Xiaohua Chen, Xiaoshuai Hao, Hehan Li, Andi Zhang, Wenxuan Wang, Kaiyan Zhang, Guoli Jia, Lingling Li, Zhiwu Lu, Yang Lu, Yike Guo

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Xiamen University(厦门大学) The Hong Kong University of Science and Technology(香港理工大学) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10342 2025-10-14 cs.CV 79%

Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis

Yu-Hsuan Lin

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 7 pages, 4 figures. Preprint submitted to arXiv in October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10117 2025-10-14 cs.AI 79%

DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay

Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu, Baixuan Xu, Yauwai Yim, Chunkit Chan, Jiaxin Bai, Yangqiu Song

机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(计算机科学与工程系,香港科技大学,香港特别行政区,中国)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments EMNLP 2025 Wordplay (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10104 2025-10-14 cs.CV 79%

Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models

Minbin Huang, Runhui Huang, Chuanyang Zheng, Jingyao Li, Guoxuan Chen, Han Shi, Hong Cheng

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09358 2025-10-13 cs.CV 79%

Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models

Qihang Ma, Shengyu Li, Jie Tang, Dingkang Yang, Shaodong Chen, Yingyi Zhang, Chao Feng, Jiao Ran

机构 * ByteDance Douyin Content Group(字节跳动抖音内容团队)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments EMNLP2025. Code is avaible at https://github.com/bytedance/DynamicCoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06092 2025-10-10 cs.CV 79%

Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation

Yachun Mi, Yu Li, Yanting Li, Chen Hui, Tong Zhang, Zhixuan Li, Chenyue Song, Wei Yang Bryan Lim, Shaohui Liu

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18743 2025-10-07 cs.CV 79%

SAR-TEXT: A Large-Scale SAR Image-Text Dataset Built with SAR-Narrator and A Progressive Learning Strategy for Downstream Tasks

Yiguo He, Xinjun Cheng, Junjie Zhu, Chunping Qiu, Jun Wang, Xichuan Zhang, Qiangjuan Huang, Ke Yang

机构 * Intelligent Game and Decision Lab(智能游戏与决策实验室)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments IEEE Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25791 2025-10-01 cs.CV 79%

EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks

Yuan Gao, Sangwook Kim, Chris McIntosh

机构 * Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络) Department of Medical Biophysics, UofT(医学生物物理学系) Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络) Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学) Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute) Department of Medical Imaging, UofT(医学影像学系) Vector Institute, Toronto(向量研究所)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments MICCAI 2025

Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. MICCAI 2025. Lecture Notes in Computer Science, vol 15964. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18839 2025-09-24 cs.CV 79%

Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography

Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza

机构 * Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14958 2025-09-23 cs.CV 79%

Seeing 3D Through 2D Lenses: 3D Few-Shot Class-Incremental Learning via Cross-Modal Geometric Rectification

Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He

机构 * South China University of Technology(南方科技大学) State Key Laboratory of Subtropical Building Science(亚热带建筑科学国家重点实验室) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室) Singapore Management University(新加坡国立大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16600 2025-09-19 cs.CV 79%

Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States

Qizao Wang, Xuelin Qian, Bin Li, Yanwei Fu, Xiangyang Xue

机构 * School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学) School of Computer Science, Shanghai Key Lab of Intelligent Information Processing, Fudan University(计算机学院,上海智能信息处理重点实验室,复旦大学) Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted by TIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08679 2025-09-17 cs.CV 79%

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

Rajarshi Roy, Devleena Das, Ankesh Banerjee, Arjya Bhattacharjee, Kousik Dasgupta, Subarna Tripathi

机构 * Kalyani Government Engineering College(卡利尼政府工程学院) Intel Labs(英特尔实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏