arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-28 至 2025-10-28 共收录 129 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 16 篇

2510.23184 2025-10-28 cs.CV 88%

Finding 3D Scene Analogies with Multimodal Foundation Models

Junho Kim, Young Min Kim

机构 * Institute of New Media and Communications(新媒体与通讯研究所) Dept. of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments Accepted to FM4RoboPlan workshop at RSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15510 2025-10-28 cs.CV cs.CL 84%

Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought

Zihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang, Weiyun Wang, Hao Fei, Yidong Wang, Alex Jinpeng Wang, Zhi Chen, Wanxiang Che, Libo Qin

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology(哈尔滨工业大学社会计算与交互机器人研究中心) Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院计算与智能研究所) Text Computing and Cognitive Intelligence Ministry of Education Engineering Research Center, Guizhou University(贵州大学文字计算与认知智能教育部工程研究中心) Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室) National University of Singapore(新加坡国立大学) Peking University(北京大学) ByteDance Seed (China)(字节跳动种子(中国))

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at NeurIPS 2025;

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12116 2025-10-28 cs.CL cs.AI cs.CV 82%

Unsupervised Document and Template Clustering using Multimodal Embeddings

Phillipe R. Sampaio, Helene Maxcici

机构 * BNP Paribas Cardif Nanterre(BNP巴黎银行卡迪夫纳特尔尔分校) BNP Paribas Paris(BNP巴黎银行巴黎)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 24 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22838 2025-10-28 cs.CV 74%

Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models

Aya Nakayama, Brian Wong, Yuji Nishimura, Kaito Tanaka

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20625 2025-10-28 cs.CV 74%

T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting

Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, Michael P. Pound

机构 * University of Nottingham(诺丁汉大学) University of St Andrews(圣安德鲁大学) City University of Hong Kong(香港城市大学) Nanyang Technology University(南洋理工大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :cross-modal(title);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21879 2025-10-28 cs.CV cs.AI 73%

TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge

Shu-Hao Zhang, Wei-Cheng Tang, Chen Wu, Peng Hu, Nan Li, Liang-Jie Zhang, Qi Zhang, Shao-Qun Zhang

机构 * State Key Laboratory of Novel Software Technology, Nanjing University(新型软件技术国家重点实验室) Microsoft AI(微软人工智能)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22785 2025-10-28 cs.CV 70%

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari, Mingkun Xu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Agency for Science, Technology and Research(科技研究局) School of Information Technology(信息技术学院)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07019 2025-10-28 cs.CV 70%

A Vision-Language Foundation Model for Leaf Disease Identification

Khang Nguyen Quoc, Lan Le Thi Thu, Luyl-Da Quach

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11593 2025-10-28 cs.CV cs.AI cs.CL cs.DB cs.LG 67%

Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion

Luigi Celona, Simone Bianco, Marco Donzella, Paolo Napoletano

机构 * Department of Informatics, Systems and Communication(信息学、系统与通信系)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments This manuscript has been accepted for publication in Springer Neural Computing and Applications

Journal ref Neural Computer & Application 37, 27279-27299 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22370 2025-10-28 cs.RO cs.AI cs.CV cs.LG cs.SE 62%

BLIP-FusePPO: A Vision-Language Deep Reinforcement Learning Framework for Lane Keeping in Autonomous Vehicles

Seyed Ahmad Hosseini Miangoleh, Amin Jalal Aghdasian, Farzaneh Abdollahi

机构 * Department of Electrical Engineering, Amirkabir University of Technology (Tehran Polytechnic)(电气工程系,阿米尔卡比尔技术大学(德黑兰理工大学))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments https://github.com/Amin-A96/BLIP-FusePPO-A-Vision-Language-Deep-Reinforcement-Learning-Framework-for-Lane-Keeping-in-Autonomous.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22045 2025-10-28 cs.CV cs.AI 62%

VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT

Hyeonsu Kang, Emily Bao, Anjan Goswami

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle - Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21806 2025-10-28 cs.CV cs.AI 62%

Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval

Jiaao Yu, Mingjie Han, Tao Gong, Jian Zhang, Man Lan

机构 * School of Computer Science and Technology, East China Normal University, China(上海师范大学计算机科学与技术学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19333 2025-10-28 cs.CV 57%

A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP

Ying Dai, Wei Yu Chen

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05020 2025-10-28 cs.RO cs.AI 57%

Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System

Haokun Liu, Zhaoqi Ma, Yunong Li, Junichiro Sugihara, Yicheng Chen, Jinjie Li, Moju Zhao

机构 * DRAGON Lab at Department of Mechanical Engineering, The University of Tokyo(东京大学机械工程系DRAGON实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 18 pages, 10 figures

Journal ref Advanced Intelligent Systems, Oct. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03610 2025-10-28 cs.CV 57%

Learning Knowledge-based Prompts for Robust 3D Mask Presentation Attack Detection

Fangling Jiang, Qi Li, Bing Liu, Weining Wang, Caifeng Shan, Zhenan Sun, Ming-Hsuan Yang

机构 * School of Computer Science, University of South China(南方大学计算机科学学院) New Laboratory of Pattern Recognition, MAIS, CASIA(模式识别新实验室,MAIS,CASIA) School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学) Department of Computer Science and Engineering, University of California, Merced(加州大学默塞德分校计算机科学与工程系) Department of Computer Science and Engineering, Yonsei University(延世大学计算机科学与工程系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13317 2025-10-28 cs.LG 50%

Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models

Song-Lin Lv, Rui Zhu, Tong Wei, Yu-Feng Li, Lan-Zhe Guo

机构 * School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学) National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学) School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学)

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2507.16343 2025-10-28 cs.SD cs.AI eess.AS 84%

Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries

Pengfei Cai, Yan Song, Qing Gu, Nan Jiang, Haoyu Song, Ian McLoughlin

机构 * University of Science and Technology of China(科学技术大学) Singapore Institute of Technology(新加坡理工学院)

专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted by MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00358 2025-10-28 cs.SD cs.AI cs.LG eess.AS 84%

$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time

Sarthak Kumar Maharana, Saksham Singh Kushwaha, Baoming Zhang, Adrian Rodriguez, Songtao Wei, Yapeng Tian, Yunhui Guo

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.AI、eess.AS

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19549 2025-10-28 cs.CV 83%

DEEMO: De-identity Multimodal Emotion Recognition and Reasoning

Deng Li, Bohao Xing, Xin Liu, Baiqiang Xia, Bihan Wen, Heikki Kälviäinen

机构 * Lappeenranta-Lahti University of Technology LUT(拉普兰塔-拉赫蒂技术大学) Nanyang Technological University(南洋理工大学) Brno University of Technology(布拉格技术大学)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17113 2025-10-28 cs.CV cs.AI cs.CL 82%

MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation

Shoubin Yu, Yue Zhang, Ziyang Wang, Jaehong Yoon, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Findings; The first two authors contributed equally; Github link: https://github.com/Yui010206/MEXA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22455 2025-10-28 cs.SD cs.AI eess.AS 81%

Evaluating Multimodal Large Language Models on Core Music Perception Tasks

Brandon James Carone, Iran R. Roman, Pablo Ripollés

机构 * Department of Psychology, Music and Audio Research Laboratory(心理学系、音乐与音频研究实验室) Department of Electronic Engineering and Computer Science(电子工程与计算机科学系)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments Accepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04273 2025-10-28 cs.IR cs.CV cs.MM cs.SD eess.AS 67%

Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval

Junan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu, Xun Yang, Jixiang Zhu, Sanyuan Zhang, Jianfeng Dong

机构 * Zhejiang University(浙江大学) Peking University(北京大学) Zhejiang Gongshang University(浙江工商大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Science and Technology of China(中国科学技术大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21864 2025-10-28 cs.CL cs.AI 62%

DeepOmni: Towards Seamless and Smart Speech Interaction with Adaptive Modality-Specific MoE

Hang Shao, Heting Gao, Yunhang Shen, Jiawei Chen, Zuwei Long, Dong Yang, Ke Li, Xing Sun

机构 * Tencent(腾讯) Fudan University(复旦大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 15 篇

2509.04086 2025-10-28 cs.CV cs.MM 84%

TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph

Yaru Chen, Faegheh Sardari, Peiliang Zhang, Ruohao Guo, Yang Xiang, Zhenbo Li, Wenwu Wang

机构 * Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(视觉、语音和信号处理中心(CVSSP),萨里大学) School of Computer Science and Artificial Intelligence, Wuhan University of Technology(计算机科学与人工智能学院,武汉理工大学) National Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室,北京大学智能科学与技术学院) College of Information and Electrical Engineering, China Agricultural University(信息与电子工程学院,中国农业大学)

专题命中 视频多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17394 2025-10-28 cs.CV cs.AI 81%

HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs

Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin

机构 * Xinjiang University(新疆大学) Xi'an Jiaotong University(西安交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21761 2025-10-28 cs.RO cs.AI cs.CV 81%

J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino

机构 * Division of Information Science, NAIST(NAIST信息科学系) Guardian Robot Project, RIKEN(RIKEN守护机器人项目)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06499 2025-10-28 q-bio.NC cs.AI 79%

The ISLab Solution to the Algonauts Challenge 2025: A Multimodal Deep Learning Approach to Brain Response Prediction

Andrea Corsico, Giorgia Rigamonti, Simone Zini, Luigi Celona, Paolo Napoletano

机构 * Department of Informatics, Systems and Communication, University of Milano-Bicocca(信息学、系统与通信系,米兰-比科卡大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01481 2025-10-28 cs.CV cs.LG 74%

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du, Fuxiao Liu, Tianyi Zhou, Dinesh Manocha, Jordan Lee Boyd-Graber

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09732 2025-10-28 cs.LG q-bio.PE stat.AP 71%

Continental-scale habitat distribution modelling with multimodal earth observation foundation models

Sara Si-Moussi, Stephan Hennekens, Sander Mucher, Stan Los, Yoann Cartier, Borja Jiménez-Alfaro, Fabio Attorre, Jens-Christian Svenning, Wilfried Thuiller

专题命中 视频多模态 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21786 2025-10-28 cs.CV cs.AI cs.MM 67%

EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction

Qile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang, Chao Tong

机构 * Beihang University(北京航空航天大学) University of Science and Technology Beijing(北京科技大学) School of Computer Science and Engineering(计算机科学与工程学院) State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 15 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏