arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-04 至 2025-11-04 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 15 篇

2511.00427 2025-11-04 cs.CV cs.AI 84%

Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection

Daichi Zhang, Tong Zhang, Jianmin Bao, Shiming Ge, Sabine Süsstrunk

机构 * School of Computer and Communication Sciences, EPFL(瑞士联邦理工学院计算机与通信科学学院) Microsoft Research Asia(微软亚洲研究院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00095 2025-11-04 cs.CV cs.AI 81%

SpinalSAM-R1: A Vision-Language Multimodal Interactive System for Spine CT Segmentation

Jiaming Liu, Dingwei Fan, Junyong Zhao, Chunlin Li, Haipeng Si, Liang Sun

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学) Department of Orthopedics, Qilu Hospital, Shandong University(骨科部,齐鲁医院,山东大学) Key Laboratory of Qingdao in Medicine and Engineering(医学与工程青岛重点实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 2 Tables,5 Figures,16 Equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12149 2025-11-04 cs.CL cs.MM cs.SI 81%

Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models

Junjie Chen, Xuyang Liu, Subin Huang, Linfeng Zhang, Hang Yu

机构 * Anhui Polytechnic University (AHPU)(安徽工程大学) Shanghai University (SHU)(上海大学) Shanghai Jiao Tong University (SJTU)(上海交通大学) Sichuan University (SCU)(四川大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00916 2025-11-04 cs.CV 79%

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

Yan Shu, Chi Liu, Robin Chen, Derek Li, Bryan Dai

机构 * Ubiquant(乌比量化)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00821 2025-11-04 cs.CV 79%

OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models

Ruoxiang Huang, Xindian Ma, Rundong Kong, Zhen Yuan, Peng Zhang

机构 * Tianjin University(天津大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02095 2025-11-04 cs.CV cs.LG 79%

Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences

Hyojin Bahng, Caroline Chan, Fredo Durand, Phillip Isola

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01341 2025-11-04 cs.CL 79%

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow York University(约克大学) Mila – Quebec AI Institute(魁北克人工智能研究院) École de Technologie Supérieure(魁北克高等技术学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) CIFAR AI Chair(CIFAR人工智能 chair) Polytechnique Montréal(蒙特利尔理工学院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00831 2025-11-04 cs.CV cs.AI 79%

Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack

Xin Liu, Aoyang Zhou, Aoyang Zhou

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by NAACL2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01340 2025-11-04 cs.CV cs.CL 76%

$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles

Trishanu Das, Abhilash Nandy, Khush Bajaj, Deepiha S

机构 * Tredence Inc.(特伦德公司) Indian Institute of Technology Kharagpur(印度理工学院克拉格浦尔分校) Inria Paris-Rocquencourt(巴黎-罗克琴库特研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒姆研究实验室)

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL

Comments 7 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01694 2025-11-04 cs.LG cs.AI 57%

Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering

Hossein Abdi, Mingfei Sun, Wei Pan

机构 * The University of Manchester(曼彻斯特大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01550 2025-11-04 cs.AI 57%

Analyzing Sustainability Messaging in Large-Scale Corporate Social Media

Ujjwal Sharma, Stevan Rudinac, Ana Mićković, Willemijn van Dolen, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07409 2025-11-04 cs.CV cs.LG 57%

MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification

Anh-Tien Nguyen, Duy Minh Ho Nguyen, Nghiem Tuong Diep, Trung Quoc Nguyen, Nhat Ho, Jacqueline Michelle Metsch, Miriam Cindy Maurer, Daniel Sonntag, Hanibal Bohnenberger, Anne-Christin Hauschild

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Published in Transactions on Machine Learning Research (09/2025)

Journal ref Transactions on Machine Learning Research (09/2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00446 2025-11-04 cs.CV cs.CR cs.LG 57%

ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training

Xin Yao, Haiyang Zhao, Yimin Chen, Jiawei Guo, Kecheng Huang, Ming Zhao

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) Miner School of Computer & Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校Miner计算机与信息科学学院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02866 2025-11-04 cs.CV 57%

OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery

Xiucheng Liang, Jinheng Xie, Tianhong Zhao, Rudi Stouffs, Filip Biljecki

机构 * Department of Architecture, National University of Singapore(建筑系,新加坡国立大学) Department of Electrical and Computer Engineering, National University of Singapore(电气与计算机工程系,新加坡国立大学) School of Artificial Intelligence, Shenzhen Technology University(人工智能学院,深圳科技大学) Department of Real Estate, National University of Singapore(房地产系,新加坡国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing 230: 918-942, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08478 2025-11-04 cs.IR cs.AI cs.LG 57%

Federated Vision-Language-Recommendation with Personalized Fusion

Zhiwei Li, Guodong Long, Jing Jiang, Chengqi Zhang, Qiang Yang

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 15 pages, 10 figures, 7 tables, conference

详情

展开后加载摘要…

URL PDF HTML 收藏