arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4672 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4672 篇

2503.16843 2025-03-24 cs.CV 79%

LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models

Jian Liang, Wenke Huang, Guancheng Wan, Qu Yang, Mang Ye

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16413 2025-03-21 cs.CV cs.RO 79%

M3: 3D-Spatial MultiModal Memory

Xueyan Zou, Yuchen Song, Ri-Zhao Qiu, Xuanbin Peng, Jianglong Ye, Sifei Liu, Xiaolong Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICLR2025 homepage: https://m3-spatial-memory.github.io code: https://github.com/MaureenZOU/m3-spatial

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15940 2025-03-21 cs.CV 79%

UniCrossAdapter: Multimodal Adaptation of CLIP for Radiology Report Generation

Yaxiong Chen, Chuang Du, Chunlei Li, Jingliang Hu, Yilei Shi, Shengwu Xiong, Xiao Xiang Zhu, Lichao Mou

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CV

Comments MICCAI 2024 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23996 2025-03-18 cs.LG cs.AI cs.IT math.IT 79%

An Information Criterion for Controlled Disentanglement of Multimodal Data

Chenyu Wang, Sharut Gupta, Xinyi Zhang, Sana Tonekaboni, Stefanie Jegelka, Tommi Jaakkola, Caroline Uhler

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00153 2025-03-12 cs.CV cs.LG 79%

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

Kunyang Han, Yibo Hu, Mengxue Qu, Hailin Shi, Yao Zhao, Yunchao Wei

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07090 2025-03-11 stat.ML cs.AI cs.LG 79%

Generative Distribution Prediction: A Unified Approach to Multimodal Learning

Xinyu Tian, Xiaotong Shen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments 31 pages 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20548 2025-03-11 cs.RO cs.AI cs.HC 79%

Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant

Anxing Xiao, Nuwan Janaka, Tianrun Hu, Anshul Gupta, Kaixin Li, Cunjun Yu, David Hsu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00260 2025-03-11 cs.CV 79%

GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning

Jun Wang, Hao Ruan, Liangjian Wen, Yong Dai, Mingjie Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16071 2025-03-04 cs.CV 79%

DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution

Yuzhong Zhao, Feng Liu, Yue Liu, Mingxiang Liao, Chen Gong, Qixiang Ye, Fang Wan

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted in CVPR 2025. Code is available at https://github.com/callsys/DynRefer

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19450 2025-02-28 cs.CV 79%

CLIP-Optimized Multimodal Image Enhancement via ISP-CNN Fusion for Coal Mine IoVT under Uneven Illumination

Shuai Wang, Shihao Zhang, Jiaqi Wu, Zijian Tian, Wei Chen, Tongzhu Jin, Miaomiao Xue, Zehua Wang, Fei Richard Yu, Victor C. M. Leung

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16586 2025-02-25 cs.CV 79%

Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review

Pei Fu, Tongkun Guan, Zining Wang, Zhentao Guo, Chen Duan, Hao Sun, Boming Chen, Jiayao Ma, Qianyi Jiang, Kai Zhou, Junfeng Luo

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14420 2025-02-24 cs.RO cs.CV cs.LG 79%

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Zhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Ran Cheng, Yaxin Peng, Chaomin Shen, Feifei Feng

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13766 2025-02-20 cs.CL 79%

GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking

Florian Schneider, Carolin Holtermann, Chris Biemann, Anne Lauscher

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13363 2025-02-20 cs.CV cs.LG 79%

Pretrained Image-Text Models are Secretly Video Captioners

Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush Vosoughi

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV

Comments Accepted to the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL 2025). The first two authors contributed equally and were listed in random order

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11073 2025-02-18 cs.CL 79%

Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions

Ming Shan Hee, Roy Ka-Wei Lee

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Preprint. Accepted at ICWSM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11024 2025-02-18 cs.CV 79%

TPCap: Unlocking Zero-Shot Image Captioning with Trigger-Augmented and Multi-Modal Purification Modules

Ruoyu Zhang, Lulu Wang, Yi He, Tongling Pan, Zhengtao Yu, Yingna Li

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03888 2025-02-18 cs.CL 79%

Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models

Minh Duc Bui, Katharina von der Wense, Anne Lauscher

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to NAACL 2025 Main (Camera-Ready Version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10455 2025-02-18 cs.LG cs.MM 79%

E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection

Junjie Wu, Yumeng Fu, Nan Yu, Guohong Fu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10057 2025-01-20 cs.CL 79%

MSTS: A Multimodal Safety Test Suite for Vision-Language Models

Paul Röttger, Giuseppe Attanasio, Felix Friedrich, Janis Goldzycher, Alicia Parrish, Rishabh Bhardwaj, Chiara Di Bonaventura, Roman Eng, Gaia El Khoury Geagea, Sujata Goswami, Jieun Han, Dirk Hovy, Seogyeong Jeong, Paloma Jeretič, Flor Miriam Plaza-del-Arco, Donya Rooein, Patrick Schramowski, Anastassia Shaitarova, Xudong Shen, Richard Willats, Andrea Zugarini, Bertie Vidgen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05901 2025-01-14 cs.CV 79%

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Ziheng Wu, Zhenghao Chen, Ruipu Luo, Can Zhang, Yuan Gao, Zhentao He, Xian Wang, Haoran Lin, Minghui Qiu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12926 2025-01-14 cs.AI 79%

MM-PhyRLHF: Reinforcement Learning Framework for Multimodal Physics Question-Answering

Janak Kapuriya, Chhavi Kirtani, Apoorv Singh, Jay Saraf, Naman Lal, Jatin Kumar, Adarsh Raj Shivam, Astha Verma, Avinash Anand, Rajiv Ratn Shah

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03786 2025-01-08 cs.CV 79%

KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration

Chengyuan Li, Suyang Zhou, Jieping Kong, Lei Qi, Hui Xue

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18558 2025-01-07 cs.CL 79%

Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Shuhao Gu, Jialing Zhang, Siyuan Zhou, Kevin Yu, Zhaohu Xing, Liangdong Wang, Zhou Cao, Jintao Jia, Zhuoyi Zhang, Yixuan Wang, Zhenchong Hu, Bo-Wen Zhang, Jijie Li, Dong Liang, Yingli Zhao, Songjing Wang, Yulong Ao, Yiming Ju, Huanhuan Ma, Xiaotong Li, Haiwen Diao, Yufeng Cui, Xinlong Wang, Yaoqi Liu, Fangxiang Feng, Guang Liu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17800 2024-12-24 cs.CV 79%

Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection

Yitong Chen, Wenhao Yao, Lingchen Meng, Sihong Wu, Zuxuan Wu, Yu-Gang Jiang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Code is available at https://github.com/Row11n/Prova/tree/main

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07543 2024-12-23 cs.CV 79%

Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning

Chenyu Yang, Xizhou Zhu, Jinguo Zhu, Weijie Su, Junjie Wang, Xuan Dong, Wenhai Wang, Lewei Lu, Bin Li, Jie Zhou, Yu Qiao, Jifeng Dai

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15894 2024-12-23 cs.CV 79%

Craft: Cross-modal Aligned Features Improve Robustness of Prompt Tuning

Jingchen Sun, Rohan Sharma, Vishnu Suresh Lokhande, Changyou Chen

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05530 2024-12-10 cs.CV 79%

CLIP-TNseg: A Multi-Modal Hybrid Framework for Thyroid Nodule Segmentation in Ultrasound Images

Xinjie Sun, Boxiong Wei, Yalong Jiang, Liquan Mao, Qi Zhao

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 4 pages, 2 figures, submitted to IEEE Signal Processing Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.13061 2024-12-09 eess.IV cs.CV 79%

Prediction of Thrombectomy Functional Outcomes using Multimodal Data

Zeynel A. Samak, Philip Clatworthy, Majid Mirmehdi

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at Medical Image Understanding and Analysis (MIUA) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15459 2024-11-26 cs.CV 79%

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking

Xinqi Liu, Li Zhou, Zikun Zhou, Jianqiu Chen, Zhenyu He

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14402 2024-11-22 cs.CV cs.LG 79%

Multimodal Autoregressive Pre-training of Large Vision Encoders

Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor Guilherme Turrisi da Costa, Louis Béthune, Zhe Gan, Alexander T Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, Alaaeldin El-Nouby

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments https://github.com/apple/ml-aim

详情

展开后加载摘要…

URL PDF HTML 收藏