arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-03 至 2025-09-03 共收录 120 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 15 篇

2509.00053 2025-09-03 cs.MM cs.AI cs.CL 90%

Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?

Shuo Liu, Di Yao, Yan Lin, Gao Cong, Jingping Bi

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Department of Computer Science, Aalborg University(奥胡斯大学计算机科学系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(title,abstract);image-text(abstract);分类 cs.CL、cs.AI、cs.MM

Comments 20 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09623 2025-09-03 cs.CV cs.AI cs.CL 87%

Cross-Modal Adapter for Vision-Language Retrieval

Haojun Jiang, Jianke Zhang, Rui Huang, Chunjiang Ge, Zanlin Ni, Shiji Song, Gao Huang

机构 * Department of Automation, BNRist, Tsinghua University(自动化系,北京理工大学,清华大学) School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)

专题命中 图文多模态 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The accepted manuscript by Pattern Recognition 25 Journal. The published journal article is available at: https://doi.org/10.1016/j.patcog.2024.111144

Journal ref Pattern Recognition 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01644 2025-09-03 cs.CV 83%

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

Yanqing Liu, Xianhang Li, Letian Zhang, Zirui Wang, Zeyu Zheng, Yuyin Zhou, Cihang Xie

机构 * University of California Santa Cruz(加州大学圣克鲁兹分校) Apple(苹果公司) University of California Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00039 2025-09-03 cs.CV 83%

AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models

Yuqi Li, Chuanguang Yang, Junhao Dong, Zhengtao Yao, Haoyan Xu, Zeyu Dong, Hansheng Zeng, Zhulin An, Yingli Tian

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02129 2025-09-03 cs.LG cs.CV 79%

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

Jintao Cheng, Weibin Li, Jiehao Luo, Xiaoyu Tang, Zhijian He, Jin Wu, Yao Zou, Wei Zhang

机构 * Hong Kong University of Science(香港科学与技术大学) South China Normal University, Shanwei, Guangdong, China(华南师范大学,汕尾,广东,中国) Shenzhen Technology University, Shenzhen, Guangdong, China(深圳科技大学,深圳,广东,中国) University of Science(科学大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20860 2025-09-03 cs.CV 79%

FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models

Mainak Singha, Subhankar Roy, Sarthak Mehrotra, Ankit Jha, Moloud Abdar, Biplab Banerjee, Elisa Ricci

机构 * University of Trento(特伦托大学) University of Bergamo(贝拉姆奥大学) Indian Institute of Technology Bombay(印度班加罗尔理工学院) LNMIIT Jaipur(斋普尔LNMIIT) The University of Queensland(昆士兰大学) Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16044 2025-09-03 cs.CV 79%

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, Jianwei Yin

机构 * Zhejiang University(浙江大学) Om AI Research(Om AI 研究所) Binjiang Institute of Zhejiang University(浙江大学滨江研究院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by EMNLP-2025 Main. Project page: https://szhanz.github.io/zoomeye/

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00449 2025-09-03 cond-mat.mtrl-sci cond-mat.dis-nn cs.LG 78%

Molecular Identification from AFM images using the IUPAC Nomenclature and Attribute Multimodal Recurrent Neural Networks

Jaime Carracedo-Cosme, Carlos Romero-Muñiz, Pablo Pou, Rubén Pérez

专题命中 图文多模态 :multimodal(title,abstract)

Comments 30 pages, 4 figures, 2 tables, includes supplementary information (with additional 21 pages, 9 figures, 1 table)

Journal ref ACS Appl. Mater. Interfaces 15, 22692-22704 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16680 2025-09-03 cs.CV 77%

AeroReformer: Aerial Referring Transformer for UAV-based Referring Image Segmentation

Rui Li, Xiaowei Zhao

机构 * Intelligent Control \& Smart Energy (ICSE) Research Group, School of Engineering, University of Warwick, Coventry, CV4 7AL, UK

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00284 2025-09-03 cs.CV cs.AI 73%

Generative AI for Industrial Contour Detection: A Language-Guided Vision System

Liang Gong, Tommy, Wang, Sara Chaker, Yanchen Dong, Fouad Bousetouane, Brenden Morton, Mark Mendez

机构 * The University of Chicago(芝加哥大学) FabTrack

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17243 2025-09-03 cs.CV cs.AI cs.CL 67%

CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models

Zicong Tang, Ziyang Ma, Suqing Wang, Zuchao Li, Lefei Zhang, Hai Zhao, Yun Li, Qianren Wang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院) Cognitive AI Lab, Shanghai Huawei Technologies, China(上海华为技术有限公司认知人工智能实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16647 2025-09-03 cs.CV cs.AI 62%

Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models

Sushant Gautam, Michael A. Riegler, Pål Halvorsen

机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet), Norway(Simula数字工程中心(SimulaMet)) Oslo Metropolitan University (OsloMet), Norway(奥斯陆 Metropolitan 大学(OsloMet)) Simula Research Laboratory, Norway(Simula研究实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted as a full paper at the 38th IEEE International Symposium on Computer-Based Medical Systems (CBMS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06794 2025-09-03 cs.CV cs.CL 62%

Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study

Yizheng Sun, Hao Li, Chang Xu, Hongpeng Zhou, Chenghua Lin, Riza Batista-Navarro, Jingyuan Sun

机构 * University of Manchester(曼彻斯特大学) Microsoft Research(微软研究院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13859 2025-09-03 cs.CV 57%

Learning Visual Proxy for Compositional Zero-Shot Learning

Shiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing, Lei Zhou, Wenjun Wang

机构 * Tianjin University(天津大学) Zhejiang University(浙江大学) Zhejiang University of Technology(浙江工业大学) Hainan University(海南大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00676 2025-09-03 cs.CV cs.LG 57%

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Xiyao Wang, Chunyuan Li, Jianwei Yang, Kai Zhang, Bo Liu, Tianyi Xiong, Furong Huang

机构 * University of Maryland College Park(马里兰大学 College Park 分校) The Ohio State University(俄亥俄州立大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2509.00723 2025-09-03 cs.AI cs.MM 84%

OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination

Junzhe Chen, Tianshu Zhang, Shiyu Huang, Yuwei Niu, Chao Sun, Rongzhou Zhang, Guanyu Zhou, Lijie Wen, Xuming Hu

机构 * Tsinghua University(清华大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) OpenRL Chongqing University(重庆大学)

专题命中 音频语音多模态 :omni-modal(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03084 2025-09-03 cs.LG cs.AI 79%

Adversarial Attacks in Multimodal Systems: A Practitioner's Survey

Shashank Kapoor, Sanjay Surendranath Girija, Lakshit Arora, Dipen Pradhan, Ankit Shetgaonkar, Aman Raj

机构 * Google(谷歌)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted in IEEE COMPSAC 2025

Journal ref 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01246 2025-09-03 cs.HC cs.RO 50%

An AI-Based Shopping Assistant System to Support the Visually Impaired

Larissa R. de S. Shibata, Ankit A. Ravankar, Jose Victorio Salazar Luces, Yasuhisa Hirata

机构 * Department of Robotics, Tohoku University(机器人系,东京东京大学)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 7 pages, Accepted for 2025 SICE-FES conference (IEEE)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 13 篇

2509.01177 2025-09-03 cs.CV cs.AI cs.HC eess.SP 81%

DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion

Junxiang Liu, Junming Lin, Jiangtong Li, Jie Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00357 2025-09-03 cs.CV cs.AI cs.LG 81%

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu, Danny T. M. Chan, Nassir Navab, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * Hong Kong Institute of Science & Innovation(香港科学与工业创新研究院) CAMP, Technische Universität München(CAMP,慕尼黑技术大学) Department of Surgery, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院外科部)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01338 2025-09-03 cs.AI 79%

Conformal Predictive Monitoring for Multi-Modal Scenarios

Francesca Cairoli, Luca Bortolussi, Jyotirmoy V. Deshmukh, Lars Lindemann, Nicola Paoletti

机构 * University of Trieste, Trieste, Italy(特里埃斯特大学) University of Southern California, Los Angeles, California(南加州大学) King's College London, London, United Kingdom(伦敦国王学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04817 2025-09-03 cs.CV 79%

LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Hongjie Zhang, Lu Dong, Yi Liu, Yifei Huang, Yali Wang, Limin Wang, Yu Qiao

机构 * OpenGVLab, Shanghai AI Laboratory, China(OpenGVLab,上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Honor Device Co.,Ltd(荣耀设备有限公司) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) Nanjing University(南京大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15825 2025-09-03 cs.CL q-fin.ST 79%

Enhancing Cryptocurrency Sentiment Analysis with Multimodal Features

Chenghao Liu, Aniket Mahanti, Ranesh Naha, Guanghao Wang, Erwann Sbai

机构 * Department of Computer Science, The University of Auckland(计算机科学系,奥克兰大学) School of Information Systems, Queensland University of Technology(信息系统学院,昆士兰技术大学) Department of Economics, The University of Auckland(经济学系,奥克兰大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01591 2025-09-03 cs.CV 79%

Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval

Adriano Fragomeni, Dima Damen, Michael Wray

机构 * School of Computer Science University of Bristol(计算机科学学院英国布里斯托尔大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05395 2025-09-03 cs.CV cs.IR cs.MM eess.IV 73%

TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations

Mert Can Cakmak, Nitin Agarwal, Diwash Poudel

机构 * Computer and Information Science, University of Arkansas - Little Rock(计算机与信息科学,亚拉荷马州立大学) ICSI, University of California, Berkeley(ICSI,加州大学伯克利分校) COSMOS Research Center, University of Arkansas - Little Rock(COSMOS研究中心,亚拉荷马州立大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13919 2025-09-03 cs.CV cs.AI cs.CL cs.LG cs.RO 67%

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang, Orr Zohar, Zeyu Wang, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21496 2025-09-03 cs.CV cs.AI 62%

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, Xuanyu Zheng, Yepeng Tang, Dahua Lin, Lewei Lu

机构 * Sensetime(秒氏科技)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01383 2025-09-03 cs.CV cs.MM 62%

Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning

Long Zhang, Peipei Song, Jianfeng Dong, Kun Li, Xun Yang

机构 * University of Science and Technology of China(中国科学技术大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发式智能感知与认知实验室) Zhejiang Gongshang University(浙江工商大学) ReLER, CCAI, Zhejiang University(ReLER,中国计算机学会,浙江大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00210 2025-09-03 cs.CV cs.AI 62%

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

Jinzhou Tang, Jusheng zhang, Sidi Liu, Waikit Xiu, Qinhan Lv, Xiying Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16834 2025-09-03 cs.LG cs.AI physics.ao-ph 57%

Improving Significant Wave Height Prediction Using Chronos Models

Yilin Zhai, Hongyuan Shi, Chao Zhan, Qing Wang, Zaijin You, Nan Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2403.07815 by other authors

Journal ref Ocean Engineering, Volume 341, Part 2, 1 December 2025, Article 122502

详情

展开后加载摘要…

URL PDF HTML 收藏