arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4682 篇

2508.03654 2025-08-06 cs.CL cs.CV 81%

Can Large Vision-Language Models Understand Multimodal Sarcasm?

Xinyu Wang, Yue Zhang, Liqiang Jing

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18937 2025-08-05 cs.CV cs.CL 81%

Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

Mahmoud Ahmed, Junjie Fei, Jian Ding, Eslam Mohamed Bakr, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(卡斯特大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09256 2025-07-15 cs.CV cs.IR cs.MM 81%

Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching

Junyu Chen, Yihua Gao, Mingyuan Ge, Mingyong Li

机构 * College of Computer and Information Science, Chongqing Normal University(重庆师范大学计算机与信息科学学院) School of Big Data and Software Engineering, Chongqing University(重庆大学大数据与软件工程学院)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by the Knowledge-Based Systems(KBS), 2025

Journal ref Volume 316, 12 May 2025, 113355

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06406 2025-06-26 cs.CL cs.AI 81%

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities

Guoyang Xia, Yifeng Ding, Fengfa Li, Lei Ren, Wei Chen, Fangxiang Feng, Xiaojie Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13038 2025-06-18 cs.CV cs.MM 81%

HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs

Zijian Zhang, Xuecheng Wu, Danlei Huang, Siyu Yan, Chong Peng, Xuezhi Cao

机构 * Xi'an Jiaotong University(西安交通大学) East China Normal University(华东师范大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01208 2025-06-17 cs.CV cs.CL 81%

Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models

Tianjie Ju, Yi Hua, Hao Fei, Zhenyu Shao, Yubin Zheng, Haodong Zhao, Mong-Li Lee, Wynne Hsu, Zhuosheng Zhang, Gongshen Liu

机构 * Shanghai Jiao Tong University(上海交通大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13509 2025-06-17 cs.CL cs.AI cs.LG 81%

ProMedTS: A Self-Supervised, Prompt-Guided Multimodal Approach for Integrating Medical Text and Time Series

Shuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, Wei Bi, Yida Xu, Guo Li, Xian Yang

机构 * Hong Kong Baptist University(香港 Baptist 大学) Shanxi University(山西大学) Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海研究院) Tencent AI Lab(腾讯人工智能实验室) Manchester Metropolitan University(曼彻斯特 Metropolitan 大学) The University of Manchester(曼彻斯特大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments This paper is accepted by ACL2025(Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10913 2025-06-11 cs.SD cs.AI eess.AS 81%

Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning

Choi Changin, Lim Sungjun, Rhee Wonjong

机构 * Interdisciplinary Program in Artificial Intelligence(人工智能交叉学科项目) Department of Intelligence and Information(智能与信息系)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13427 2025-06-06 cs.AI cs.CV 81%

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision

Lingxiao Du, Fanqing Meng, Zongkai Liu, Zhixiang Zhou, Ping Luo, Qiaosheng Zhang, Wenqi Shao

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01419 2025-06-05 cs.CV cs.AI 81%

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Mingi Jung, Saehyung Lee, Eunji Kim, Sungroh Yoon

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20753 2025-05-28 cs.CV cs.AI 81%

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

Yufei Zhan, Hongyin Zhao, Yousong Zhu, Shurong Zheng, Fan Yang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14846 2025-05-22 cs.CV cs.CL 81%

Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

Yue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta, Luca Weihs, Andrew Head, Mark Yatskar, Chris Callison-Burch, Ranjay Krishna, Aniruddha Kembhavi, Christopher Clark

机构 * University of Pennsylvania(宾夕法尼亚大学) Allen Institute for Artificial Intelligence(人工智能研究所)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Published in ACL 2025, project page: https://yueyang1996.github.io/cosyn/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14627 2025-05-21 cs.AI cs.CL 81%

Debating for Better Reasoning: An Unsupervised Multimodal Approach

Ashutosh Adhikari, Mirella Lapata

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10526 2025-05-20 cs.LG cs.CL cs.CV 81%

MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models

Mugilan Ganesan, Shane Segal, Ankur Aggarwal, Nish Sinnadurai, Sean Lie, Vithursan Thangarasa

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Main paper: 11 pages, 4 figures, 3 tables. Supplementary: 1 page

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22330 2025-05-08 cs.CV cs.CL cs.LG 81%

Vision-Language Models Create Cross-Modal Task Representations

Grace Luo, Trevor Darrell, Amir Bar

机构 * University of California, Berkeley, USA(加州大学伯克利分校)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19443 2025-04-29 cs.CV cs.AI 81%

CLIP-KOA: Enhancing Knee Osteoarthritis Diagnosis with Multi-Modal Learning and Symmetry-Aware Loss Functions

Yejin Jeong, Donghun Lee

机构 * Department of Mathematics, Korea University(数学系,韩国大学)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19127 2025-04-29 cs.CV cs.MM 81%

DeepSPG: Exploring Deep Semantic Prior Guidance for Low-light Image Enhancement with Multimodal Learning

Jialang Lu, Huayu Zhao, Huiyu Zhai, Xingxing Yang, Shini Han

机构 * School of Cyber Science and Technology, Hubei University(湖北大学计算机科学与技术学院) Department of Electrical Automation Design, Beijing Shougang International Engineering Technology(北京首钢国际工程技术部) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学部) School of Computer Science and Technology, Harbin University of Science and Technology(哈尔滨理工大学计算机科学与技术学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ICMR 2025 Main track. Code is available at https://github.com/Wenyuzhy/DeepSPG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17826 2025-04-28 cs.CV cs.AI 81%

FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model

Kaicheng Pang, Xingxing Zou, Waikeung Wong

机构 * Laboratory for Artificial Intelligence in Design(人工智能设计实验室) School of Fashion and Textiles(时尚与纺织学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18203 2025-04-24 cs.CV cs.CL 81%

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Di Zhang, Junxian Li, Jingdi Lei, Xunzhi Wang, Yujie Liu, Zonglin Yang, Jiatong Li, Weida Wang, Suorong Yang, Jianbo Wu, Peng Ye, Wanli Ouyang, Dongzhan Zhou

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiaotong University(上海交通大学) Nankai University(南开大学) Shanghai University(上海大学) Nanyang Technological University(南洋理工大学) Hong Kong Polytechnic University(香港理工大学) Tongji University(同济大学) Nanjing University(南京大学) University of California, Merced(加州大学梅德福分校) Chinese University of Hong Kong(香港中文大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 16 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13351 2025-04-21 cs.RO cs.AI cs.HC cs.LG cs.MM 81%

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

Chen Wang, Fei Xia, Wenhao Yu, Tingnan Zhang, Ruohan Zhang, C. Karen Liu, Li Fei-Fei, Jie Tan, Jacky Liang

机构 * Google DeepMind(谷歌DeepMind) Stanford University(斯坦福大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07521 2025-04-18 cs.AI cs.MM 81%

Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models

Yuxiang Lin, Jingdong Sun, Zhi-Qi Cheng, Jue Wang, Haomin Liang, Zebang Cheng, Yifei Dong, Jun-Yan He, Xiaojiang Peng, Xian-Sheng Hua

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted at CVPR Workshop NEXD 2025. 21 pages, Project: https://github.com/Lum1104/EIBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12879 2025-04-17 cs.CL cs.AI cs.LG 81%

Large Visual-Language Models Are Also Good Classifiers: A Study of In-Context Multimodal Fake News Detection

Ye Jiang, Yimin Wang

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Withdraw for new experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07643 2025-04-11 cs.IR cs.CL cs.CV 81%

CollEX -- A Multimodal Agentic RAG System Enabling Interactive Exploration of Scientific Collections

Florian Schneider, Narges Baba Ahmadi, Niloufar Baba Ahmadi, Iris Vogel, Martin Semmann, Chris Biemann

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04903 2025-04-09 cs.CV cs.AI 81%

Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision

Yuandong Pu, Le Zhuo, Kaiwen Zhu, Liangbin Xie, Wenlong Zhang, Xiangyu Chen, Peng Gao, Yu Qiao, Chao Dong, Yihao Liu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05299 2025-04-08 cs.AI cs.CV 81%

SmolVLM: Redefining small and efficient multimodal models

Andrés Marafioti, Orr Zohar, Miquel Farré, Merve Noyan, Elie Bakouch, Pedro Cuenca, Cyril Zakka, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, Vaibhav Srivastav, Joshua Lochner, Hugo Larcher, Mathieu Morlon, Lewis Tunstall, Leandro von Werra, Thomas Wolf

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14538 2025-04-02 eess.IV cs.AI cs.CV cs.LG 81%

Vision-Language Models for Acute Tuberculosis Diagnosis: A Multimodal Approach Combining Imaging and Clinical Data

Ananya Ganapthy, Praveen Shastry, Naveen Kumarasami, Anandakumar D, Keerthana R, Mounigasri M, Varshinipriya M, Kishore Prasath Venkatesh, Bargava Subramanian, Kalyan Sivasailam

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14694 2025-03-20 cs.CL cs.CV 81%

HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding

Rui Yang, Lin Song, Yicheng Xiao, Runhui Huang, Yixiao Ge, Ying Shan, Hengshuang Zhao

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12446 2025-03-18 cs.CV cs.AI 81%

BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries

Tianle Li, Yongming Rao, Winston Hu, Yu Cheng

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06232 2025-03-18 cs.CL cs.CV 81%

Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning

Yanjun Chen, Yirong Sun, Xinghao Chen, Jian Wang, Xiaoyu Shen, Wenjie Li, Wei Zhang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08967 2025-03-18 cs.CV cs.AI 81%

PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning

Qifeng Zhou, Wenliang Zhong, Yuzhi Guo, Michael Xiao, Hehuan Ma, Junzhou Huang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏