arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-30 至 2025-10-30 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2506.21710 2025-10-30 cs.CV 86%

FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering

Liangyu Zhong, Fabio Rosenthal, Joachim Sicking, Fabian Hüger, Thorsten Bagdonat, Hanno Gottschalk, Leo Schwinn

机构 * Technical University of Berlin(柏林技术大学) Technical University of Munich(慕尼黑技术大学) CARIAD SE Volkswagen AG(大众集团)

专题命中 图文多模态 :MLLM(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 - main track. Project page: https://focus-mllm-vqa.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04508 2025-10-30 cs.CL 85%

Adapter-state Sharing CLIP for Parameter-efficient Multimodal Sarcasm Detection

Soumyadeep Jana, Sahil Danayak, Sanasam Ranbir Singh

机构 * Department of Computer Science(计算机科学系) Engineering, Indian Institute of Technology Guwahati, India(工程系、印度理工学院瓜哇提学院、印度)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19311 2025-10-30 cs.CV cs.AI 84%

DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment

Weizhi Chen, Yupeng Deng, Jin Wei, Jingbo Chen, Jiansheng Chen, Yuman Feng, Zhihao Xi, Diyou Liu, Kai Li, Yu Meng

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院 aerospace information research institute) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Information Network Security, People’s Public Security University of China(中国人民公安大学信息网络安全学院)

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25303 2025-10-30 cs.CL 83%

Teaching Sarcasm: Few-Shot Multimodal Sarcasm Detection via Distillation to a Parameter-Efficient Student

Soumyadeep Jana, Sanasam Ranbir Singh

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Indian Institute of Technology Guwahati(印度理工学院古瓦哈蒂)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25070 2025-10-30 cs.CV 70%

Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments

Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25179 2025-10-30 cs.AI 57%

Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University, Australia(计算机学院,麦考瑞大学,澳大利亚)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25175 2025-10-30 cs.CV 57%

Test-Time Adaptive Object Detection with Foundation Model

Yingjie Gao, Yanan Zhang, Zhi Cai, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment, Beihang University(复杂与关键软件环境国家重点实验室,北京航空航天大学) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25051 2025-10-30 cs.CV cs.LG 57%

Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models

Shunjie-Fabian Zheng, Hyeonjun Lee, Thijs Kooi, Ali Diba

机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学院第一医学部,LMU慕尼黑大学医院) Lunit Inc.(Lunit公司)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏