arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-03 至 2025-09-03 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 15 篇

2509.00053 2025-09-03 cs.MM cs.AI cs.CL 90%

Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?

Shuo Liu, Di Yao, Yan Lin, Gao Cong, Jingping Bi

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Department of Computer Science, Aalborg University(奥胡斯大学计算机科学系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(title,abstract);image-text(abstract);分类 cs.CL、cs.AI、cs.MM

Comments 20 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09623 2025-09-03 cs.CV cs.AI cs.CL 87%

Cross-Modal Adapter for Vision-Language Retrieval

Haojun Jiang, Jianke Zhang, Rui Huang, Chunjiang Ge, Zanlin Ni, Shiji Song, Gao Huang

机构 * Department of Automation, BNRist, Tsinghua University(自动化系,北京理工大学,清华大学) School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)

专题命中 图文多模态 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The accepted manuscript by Pattern Recognition 25 Journal. The published journal article is available at: https://doi.org/10.1016/j.patcog.2024.111144

Journal ref Pattern Recognition 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01644 2025-09-03 cs.CV 83%

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

Yanqing Liu, Xianhang Li, Letian Zhang, Zirui Wang, Zeyu Zheng, Yuyin Zhou, Cihang Xie

机构 * University of California Santa Cruz(加州大学圣克鲁兹分校) Apple(苹果公司) University of California Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00039 2025-09-03 cs.CV 83%

AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models

Yuqi Li, Chuanguang Yang, Junhao Dong, Zhengtao Yao, Haoyan Xu, Zeyu Dong, Hansheng Zeng, Zhulin An, Yingli Tian

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02129 2025-09-03 cs.LG cs.CV 79%

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

Jintao Cheng, Weibin Li, Jiehao Luo, Xiaoyu Tang, Zhijian He, Jin Wu, Yao Zou, Wei Zhang

机构 * Hong Kong University of Science(香港科学与技术大学) South China Normal University, Shanwei, Guangdong, China(华南师范大学,汕尾,广东,中国) Shenzhen Technology University, Shenzhen, Guangdong, China(深圳科技大学,深圳,广东,中国) University of Science(科学大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20860 2025-09-03 cs.CV 79%

FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models

Mainak Singha, Subhankar Roy, Sarthak Mehrotra, Ankit Jha, Moloud Abdar, Biplab Banerjee, Elisa Ricci

机构 * University of Trento(特伦托大学) University of Bergamo(贝拉姆奥大学) Indian Institute of Technology Bombay(印度班加罗尔理工学院) LNMIIT Jaipur(斋普尔LNMIIT) The University of Queensland(昆士兰大学) Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16044 2025-09-03 cs.CV 79%

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, Jianwei Yin

机构 * Zhejiang University(浙江大学) Om AI Research(Om AI 研究所) Binjiang Institute of Zhejiang University(浙江大学滨江研究院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by EMNLP-2025 Main. Project page: https://szhanz.github.io/zoomeye/

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00449 2025-09-03 cond-mat.mtrl-sci cond-mat.dis-nn cs.LG 78%

Molecular Identification from AFM images using the IUPAC Nomenclature and Attribute Multimodal Recurrent Neural Networks

Jaime Carracedo-Cosme, Carlos Romero-Muñiz, Pablo Pou, Rubén Pérez

专题命中 图文多模态 :multimodal(title,abstract)

Comments 30 pages, 4 figures, 2 tables, includes supplementary information (with additional 21 pages, 9 figures, 1 table)

Journal ref ACS Appl. Mater. Interfaces 15, 22692-22704 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16680 2025-09-03 cs.CV 77%

AeroReformer: Aerial Referring Transformer for UAV-based Referring Image Segmentation

Rui Li, Xiaowei Zhao

机构 * Intelligent Control \& Smart Energy (ICSE) Research Group, School of Engineering, University of Warwick, Coventry, CV4 7AL, UK

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00284 2025-09-03 cs.CV cs.AI 73%

Generative AI for Industrial Contour Detection: A Language-Guided Vision System

Liang Gong, Tommy, Wang, Sara Chaker, Yanchen Dong, Fouad Bousetouane, Brenden Morton, Mark Mendez

机构 * The University of Chicago(芝加哥大学) FabTrack

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17243 2025-09-03 cs.CV cs.AI cs.CL 67%

CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models

Zicong Tang, Ziyang Ma, Suqing Wang, Zuchao Li, Lefei Zhang, Hai Zhao, Yun Li, Qianren Wang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院) Cognitive AI Lab, Shanghai Huawei Technologies, China(上海华为技术有限公司认知人工智能实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16647 2025-09-03 cs.CV cs.AI 62%

Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models

Sushant Gautam, Michael A. Riegler, Pål Halvorsen

机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet), Norway(Simula数字工程中心(SimulaMet)) Oslo Metropolitan University (OsloMet), Norway(奥斯陆 Metropolitan 大学(OsloMet)) Simula Research Laboratory, Norway(Simula研究实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted as a full paper at the 38th IEEE International Symposium on Computer-Based Medical Systems (CBMS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06794 2025-09-03 cs.CV cs.CL 62%

Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study

Yizheng Sun, Hao Li, Chang Xu, Hongpeng Zhou, Chenghua Lin, Riza Batista-Navarro, Jingyuan Sun

机构 * University of Manchester(曼彻斯特大学) Microsoft Research(微软研究院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13859 2025-09-03 cs.CV 57%

Learning Visual Proxy for Compositional Zero-Shot Learning

Shiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing, Lei Zhou, Wenjun Wang

机构 * Tianjin University(天津大学) Zhejiang University(浙江大学) Zhejiang University of Technology(浙江工业大学) Hainan University(海南大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00676 2025-09-03 cs.CV cs.LG 57%

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Xiyao Wang, Chunyuan Li, Jianwei Yang, Kai Zhang, Bo Liu, Tianyi Xiong, Furong Huang

机构 * University of Maryland College Park(马里兰大学 College Park 分校) The Ohio State University(俄亥俄州立大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏