arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-24 至 2025-09-24 共收录 70 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2509.18839 2025-09-24 cs.CV 79%

Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography

Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza

机构 * Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18729 2025-09-24 cs.SD 78%

MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning

Haoqin Sun, Chenyang Lyu, Xiangyu Kong, Shiwan Zhao, Jiaming Zhou, Hui Wang, Aobo Kong, Jinghua Zhao, Longyue Wang, Weihua Luo, Kaifu Zhang, Yong Qin

机构 * TMCC, College of Computer Science, Nankai University, Tianjin, China(TMCC,计算机科学学院,南开大学,天津,中国) Alibaba International Digital Commerce(阿里巴巴国际数字商务) University of Exeter(埃克塞特大学)

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18189 2025-09-24 cs.CV cs.AI 62%

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Daxiang Dong, Mingming Zheng, Dong Xu, Bairong Zhuang, Wenyu Zhang, Chunhua Luo, Haoran Wang, Zijian Zhao, Jie Li, Yuxuan Li, Hanjun Zhong, Mengyue Liu, Jieting Chen, Shupeng Li, Lun Tian, Yaping Feng, Xin Li, Donggang Jiang, Yong Chen, Yehua Xu, Duohao Qin, Chen Feng, Dan Wang, Henghua Zhang, Jingjing Ha, Jinhui He, Yanfeng Zhai, Chengxin Zheng, Jiayi Mao, Jiacheng Chen, Ruchang Yao, Ziye Yuan, Jianmin Wu, Guangjun Xie, Dou Shen

机构 * Qianfan Team, Baidu AI Cloud(百度AI云团队)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19003 2025-09-24 cs.CV 57%

Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards

Honghao Chen, Xingzhou Lou, Xiaokun Feng, Kaiqi Huang, Xinlong Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 8 篇

2509.17537 2025-09-24 cs.CV 87%

SimToken: A Simple Baseline for Referring Audio-Visual Segmentation

Dian Jin, Yanghao Zhou, Jinxing Zhou, Jiaqi Ma, Ruohao Guo, Dan Guo

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);MLLM(abstract);cross-modal(abstract)

Comments Project page: https://github.com/DianJin-HFUT/SimToken

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18912 2025-09-24 cs.CV 85%

Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation

Yunzhe Shen, Kai Peng, Leiye Liu, Wei Ji, Jingjing Li, Miao Zhang, Yongri Piao, Huchuan Lu

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13255 2025-09-24 cs.CL cs.AI cs.IR cs.LG cs.MM 85%

Automating Steering for Safe Multimodal Large Language Models

Lyucheng Wu, Mengru Wang, Ziwen Xu, Tri Cao, Nay Oo, Bryan Hooi, Shumin Deng

机构 * Zhejiang University(浙江大学) Zhejiang University - Ant Group Joint Lab of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室) National University of Singapore, NUS-NCS Joint Lab(新加坡国立大学NUS-NCS联合实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments EMNLP 2025 Main Conference. 23 pages (8+ for main); 25 figures; 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18816 2025-09-24 cs.SD cs.CL cs.MM eess.AS 82%

Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models

Junyu Wang, Ziyang Ma, Zhengding Luo, Tianrui Wang, Meng Ge, Xiaobao Wang, Longbiao Wang

机构 * 1Laboratory of Cognitive Computing Application, College of Intelligence Computing, Tianjin University, Tianjin, China 2School of Computer Science, Shanghai Jiao Tong University, Shanghai, China 3School of Electrical \& Electronic Engineering, Nanyang Technological University, Singapore 4Huiyan Technology (Tianjin) Co., Ltd, Tianjin, China

专题命中 音频语音多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CL、cs.MM、eess.AS

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18200 2025-09-24 cs.LG cs.AI cs.CL cs.RO 81%

Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought

Yu Ti Huang

机构 * Trans-disciplinary Bachelor Degree Program National Taiwan University(台湾国立大学跨学科学士学位计划)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11976 2025-09-24 cs.SD eess.AS 79%

PoolingVQ: A VQVAE Variant for Reducing Audio Redundancy and Boosting Multi-Modal Fusion in Music Emotion Analysis

Dinghao Zou, Yicheng Gong, Xiaokang Li, Xin Cao, Sunbowen Lee

机构 * Wuhan University of Science and Technology(武汉科技大学)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18706 2025-09-24 cs.HC 78%

M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition

Jiajun He, Xiaohan Shi, Cheng-Hung Hu, Jinyi Mi, Xingfeng Li, Tomoki Toda

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted by IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04561 2025-09-24 cs.CL cs.CV 76%

OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis

Run Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu, Xiong Liu, Min Yang, Yongbin Li, Longze Chen, Jiaming Li, Lei Zhang, Xiaobo Xia, Hamid Alinejad-Rokny, Fei Huang

机构 * Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室) Shenzhen Institute of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Tongyi Laboratory(通义实验室) University of New South Wales(新南威尔士大学) National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition(脑启发智能感知与认知重点实验室)

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 13 篇

2504.08727 2025-09-24 cs.CV cs.AI cs.CY 84%

Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images

Boyang Deng, Songyou Peng, Kyle Genova, Gordon Wetzstein, Noah Snavely, Leonidas Guibas, Thomas Funkhouser

机构 * Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025, Project page: https://boyangdeng.com/visual-chronicles , second and third listed authors have equal contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13707 2025-09-24 cs.CV cs.AI 84%

EventVL: Understand Event Streams via Multimodal Large Language Model

Pengteng Li, Yunfan Lu, Pinghao Song, Wuyang Li, Huizai Yao, Hui Xiong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) KU Leuven(根特大学) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院) Carleton University(卡尔顿大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02889 2025-09-24 cs.CL cs.AI cs.CV cs.MM 83%

LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture

Xidong Wang, Dingjie Song, Shunian Chen, Junyin Chen, Zhenyang Cai, Chen Zhang, Lichao Sun, Benyou Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Lehigh University(莱斯利大学) Meituan(美团)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04130 2025-09-24 cs.CV 79%

STORM: Token-Efficient Long Video Understanding for Multimodal LLMs

Jindong Jiang, Xiuyu Li, Zhijian Liu, Muyang Li, Guo Chen, Zhiqi Li, De-An Huang, Guilin Liu, Zhiding Yu, Kurt Keutzer, Sungjin Ahn, Jan Kautz, Hongxu Yin, Yao Lu, Song Han, Wonmin Byeon

机构 * NVIDIA(英伟达) Rutgers University(罗格斯大学) UC Berkeley(加州大学伯克利分校) MIT(麻省理工学院) Nanjing University(南京大学) KAIST(韩国科学技术院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18457 2025-09-24 cs.LG 78%

GluMind: Multimodal Parallel Attention and Knowledge Retention for Robust Cross-Population Blood Glucose Forecasting

Ebrahim Farahmand, Reza Rahimi Azghan, Nooshin Taheri Chatrudi, Velarie Yaa Ansu-Baidoo, Eric Kim, Gautham Krishna Gudur, Mohit Malu, Owen Krueger, Edison Thomaz, Giulia Pedrielli, Pavan Turaga, Hassan Ghasemzadeh

机构 * Arizona State University(亚利桑那州立大学) University of Miami(迈阿密大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20201 2025-09-24 cs.CV cs.AI 73%

Injecting Explainability and Lightweight Design into Weakly Supervised Video Anomaly Detection Systems

Wen-Dong Jiang, Chih-Yung Chang, Hsiang-Chuan Chang, Ji-Yuan Chen, Diptendu Sinha Roy

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18230 2025-09-24 cs.AI cs.LG 70%

Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces

Zihan Dong, Xinyu Fan, Zixiang Tang, Yunqing Li

机构 * The University of Tokyo(东京大学) Lenovo US(联想美国) Advanced AI Technology Center(高级人工智能技术中心)

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15173 2025-09-24 cs.CV cs.AI 62%

AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection

Zhipei Xu, Xuanyu Zhang, Qing Huang, Xing Zhou, Jian Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18187 2025-09-24 cs.CV cs.AI 62%

V-SenseDrive: A Privacy-Preserving Road Video and In-Vehicle Sensor Fusion Framework for Road Safety & Driver Behaviour Modelling

Muhammad Naveed, Nazia Perwaiz, Sidra Sultana, Mohaira Ahmad, Muhammad Moazam Fraz

机构 * School of Electrical Engineering and Computer Science (SEECS), National University of Sciences and Technology (NUST)(电气工程与计算机科学学院(SEECS),国立科学与技术大学(NUST))

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18183 2025-09-24 cs.CV cs.AI 62%

VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation

Jinyue Bian, Zhaoxing Zhang, Zhengyu Liang, Shiwei Zheng, Shengtao Zhang, Rong Shen, Chen Yang, Anzhou Hou

机构 * China, Beijing, Li Auto Inc.(中国北京李自动公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19245 2025-09-24 cs.CV 57%

ConViS-Bench: Estimating Video Similarity Through Semantic Concepts

Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Yiming Wang, Elisa Ricci, Paolo Rota

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07560 2025-09-24 cs.RO cs.AI 57%

Socially Pertinent Robots in Gerontological Healthcare

Xavier Alameda-Pineda, Angus Addlesee, Daniel Hernández García, Chris Reinke, Soraya Arias, Federica Arrigoni, Alex Auternaud, Lauriane Blavette, Cigdem Beyan, Luis Gomez Camara, Ohad Cohen, Alessandro Conti, Sébastien Dacunha, Christian Dondrup, Yoav Ellinson, Francesco Ferro, Sharon Gannot, Florian Gras, Nancie Gunson, Radu Horaud, Moreno D'Incà, Imad Kimouche, Séverin Lemaignan, Oliver Lemon, Cyril Liotard, Luca Marchionni, Mordehay Moradi, Tomas Pajdla, Maribel Pino, Michal Polic, Matthieu Py, Ariel Rado, Bin Ren, Elisa Ricci, Anne-Sophie Rigaud, Paolo Rota, Marta Romeo, Nicu Sebe, Weronika Sieińska, Pinchas Tandeitnik, Francesco Tonini, Nicolas Turro, Timothée Wintz, Yanchao Yu

机构 * ERM Automatismes(ERM 自动化公司) PAL Robotics(PAL 机器人公司)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19169 2025-09-24 cs.RO 50%

MagiClaw: A Dual-Use, Vision-Based Soft Gripper for Bridging the Human Demonstration to Robotic Deployment Gap

Tianyu Wu, Xudong Han, Haoran Sun, Zishang Zhang, Bangchao Huang, Chaoyang Song, Fang Wan

机构 * Design + Learning Research Group(设计+学习研究组) Southern University of Science and Technology(南方科技大学)

专题命中 视频多模态 :multi-modal(abstract)

Comments 8 pages, 4 figures, accepted to Data@CoRL2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 6 篇

2509.02783 2025-09-24 cs.LG cs.AI physics.geo-ph 84%

The Transparent Earth: A Multimodal Foundation Model for the Earth's Subsurface

Arnab Mazumder, Javier E. Santos, Noah Hobbs, Mohamed Mehana, Daniel O'Malley

机构 * Energy and Natural Resources Security Group(能源与自然资源安全组) Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)

专题命中 跨模态检索 :multimodal(title);multimodal foundation model(title);分类 cs.AI

Comments Accepted at the Neurips 2025 AI4Science Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19203 2025-09-24 cs.CV 79%

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

机构 * Queen Mary University of London(伦敦女王大学) Samsung AI Centre(三星人工智能中心) Technical University of Iași(伊阿苏技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18284 2025-09-24 cs.CV 79%

Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction

Yi Gu, Kuniaki Saito, Jiaxin Ma

机构 * OMRON SINIC X Corporation(OMRON SINIC X公司) Nara Institute of Science and Technology(名取科学技術大學院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18807 2025-09-24 cs.IR 78%

Single-Branch Network Architectures to Close the Modality Gap in Multimodal Recommendation

Christian Ganhör, Marta Moscati, Anna Hausberger, Shah Nawaz, Markus Schedl

专题命中 跨模态检索 :multimodal(title,abstract)

Comments Accepted by ACM Transactions on Recommender Systems (TORS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18672 2025-09-24 cs.HC cs.AI 74%

NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment

Ajay Narayanan Sridhar, Fuli Qiao, Nelson Daniel Troncoso Aldas, Yanpei Shi, Mehrdad Mahdavi, Laurent Itti, Vijaykrishnan Narayanan

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Independent Researcher(独立研究者) University of Southern California(南加州大学)

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏