arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-19 至 2025-09-19 共收录 48 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2405.16600 2025-09-19 cs.CV 79%

Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States

Qizao Wang, Xuelin Qian, Bin Li, Yanwei Fu, Xiangyang Xue

机构 * School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学) School of Computer Science, Shanghai Key Lab of Intelligent Information Processing, Fudan University(计算机学院,上海智能信息处理重点实验室,复旦大学) Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted by TIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15217 2025-09-19 cs.AI cs.CV cs.LG 73%

Generalizable Geometric Image Caption Synthesis

Yue Xin, Wenyuan Wang, Rui Pan, Ruida Wang, Howard Meng, Renjie Pi, Shizhe Diao, Tong Zhang

机构 * University of Illinois Urbana-Champaign (UIUC)(伊利诺伊大学厄巴纳-香槟分校) Shanghai Jiao Tong University(上海交通大学) Rutgers University(罗格斯大学) NVIDIA

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15226 2025-09-19 cs.CV 70%

Calibration-Aware Prompt Learning for Medical Vision-Language Models

Abhishek Basu, Fahad Shamshad, Ashshak Sharifdeen, Karthik Nandakumar, Muhammad Haris Khan

机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(计算机视觉系,Mohamed bin Zayed人工智能大学) Department of Computer Science and Engineering, Michigan State University (MSU)(计算机科学与工程系,密歇根州立大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted in BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14738 2025-09-19 cs.CL 70%

UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets

Pengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang, Wei Huang, Zhen Ye, Zhaowei Li, Botian Jiang, Dong Zhang, Xipeng Qiu

机构 * Fudan University(复旦大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted by Findings of EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14571 2025-09-19 cs.HC cs.AI cs.LG 57%

VisMoDAl: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models

Huanchen Wang, Wencheng Zhang, Zhiqiang Wang, Zhicong Lu, Yuxin Ma

机构 * Southern University of Science and Technology(南方科技大学) City University of Hong Kong(香港城市大学) George Mason University(乔治·马歇尔大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

Comments 11 pages, 7 figures, 1 table, accepted to IEEE VIS 2025 (IEEE Transactions on Visualization and Computer Graphics)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06937 2025-09-19 cs.RO 50%

Handle Object Navigation as Weighted Traveling Repairman Problem

Ruimeng Liu, Xinhang Xu, Shenghai Yuan, Lihua Xie

机构 * Centre for Advanced Robotics Technology Innovation (CARTIN), School of Electrical and Electronic Engineering, Nanyang Technological University(先进机器人技术创新中心(CARTIN)、电子与电气工程学院、南洋理工大学)

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 8 篇

2509.14592 2025-09-19 cs.MM cs.SD 86%

MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion

Junbo Wang, Yan Zhao, Shuo Li, Shibo Wang, Shigang Wang, Jian Wei

机构 * College of Communication Engineering, Jilin University(吉林大学通信工程学院)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14930 2025-09-19 cs.CL cs.AI 81%

Cross-Modal Knowledge Distillation for Speech Large Language Models

Enzhi Wang, Qicheng Li, Zhiyuan Tang, Yuhang Jia

机构 * TMCC, College of Computer Science, Nankai University, Tianjin, China(TMCC,计算机科学学院,南开大学,天津,中国) Tencent Ethereal Audio Lab, Tencent Corporation, Shenzhen, China(腾讯虚实音频实验室,腾讯公司,深圳,中国)

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14627 2025-09-19 cs.HC cs.AI cs.CL 81%

Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech

Taesoo Kim, Yongsik Jo, Hyunmin Song, Taehwan Kim

机构 * Artificial Intelligence Graduate School(人工智能研究生院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Published in Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14891 2025-09-19 cs.MM cs.IR cs.SD 79%

Music4All A+A: A Multimodal Dataset for Music Information Retrieval Tasks

Jonas Geiger, Marta Moscati, Shah Nawaz, Markus Schedl

机构 * Johannes Kepler University Linz(约翰内斯·开普勒大学林茨) Human-centered AI Group, AI Lab, Linz Institute of Technology(以人为本的人工智能小组、人工智能实验室、林茨技术研究所)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments 7 pages, 6 tables, IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14379 2025-09-19 eess.AS cs.LG 79%

Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior

Yochai Yemini, Rami Ben-Ari, Sharon Gannot, Ethan Fetaya

机构 * Faculty of Engineering, Bar-Ilan University(巴伊兰大学工程学院) OriginAI

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00033 2025-09-19 cs.CV cs.AI 76%

Deep Learning-Driven Multimodal Detection and Movement Analysis of Objects in Culinary

Tahoshin Alam Ishat, Mohammad Abdul Qayum

机构 * Electrical and Computer Engineering(电气与计算机工程) North South University(北南大学) Computer Engineering North South University Dhaka, Bangladesh(北南大学计算机工程系)

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16724 2025-09-19 cs.SD eess.AS 70%

SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing

Jinbo Hu, Yin Cao, Ming Wu, Zhenbo Luo, Jun Yang

机构 * Institute of Acoustics, Chinese Academy of Sciences, China(中国科学院声学研究所) MiLM Plus, Xiaomi Inc., China(小米公司) Xi’an Jiaotong Liverpool University, China(西安交通大学利物浦大学) University of Chinese Academy of Sciences, China(中国科学院大学)

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12275 2025-09-19 cs.SD cs.AI eess.AS 62%

Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering

Jinghua Zhao, Hang Su, Lichun Fan, Zhenbo Luo, Hui Wang, Haoqin Sun, Yong Qin

机构 * TMCC, College of Computer Science, Nankai University, Tianjin, China(TMCC,计算机科学学院,南开大学,天津,中国) MiLM Plus, Xiaomi Inc.(MiLM Plus,小米公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 5 pages, 1 figure, 2 tables submitted to icassp, under prereview

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2509.15178 2025-09-19 cs.CV 83%

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

Zaiquan Yang, Yuhao Liu, Gerhard Hancke, Rynson W. H. Lau

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Journal ref NeurIPS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14893 2025-09-19 cs.SD eess.AS 83%

Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic event Classification

Yuanjian Chen, Yang Xiao, Jinjie Huang

机构 * Harbin University of Science(哈尔滨理工大学) Technology The University of Melbourne(技术墨尔本大学)

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15224 2025-09-19 cs.CV 79%

Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation

Luca Bartolomei, Enrico Mannocci, Fabio Tosi, Matteo Poggi, Stefano Mattoccia

机构 * Advanced Research Center on Electronic System (ARCES)(电子系统先进研究中心) Department of Computer Science and Engineering (DISI)(计算机科学与工程系) University of Bologna(博洛尼亚大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments ICCV 2025. Code: https://github.com/bartn8/depthanyevent/ Project Page: https://bartn8.github.io/depthanyevent/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11460 2025-09-19 q-bio.QM stat.AP 50%

Mechanistic inference of stochastic gene expression from structured single-cell data

Christopher E. Miles

专题命中 视频多模态 :multimodal(abstract)

Comments submitted invited review for the `Identifiability, estimation and uncertainty in mathematical modelling' issue in Current Opinion in Systems Biology

Journal ref Curr. Opin. Syst. Biol. 42, 100555 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2509.15211 2025-09-19 cs.CL 79%

What's the Best Way to Retrieve Slides? A Comparative Study of Multimodal, Caption-Based, and Hybrid Retrieval Techniques

Petros Stylianos Giouroukis, Dimitris Dimitriadis, Dimitrios Papadopoulos, Zhenwen Shao, Grigorios Tsoumakas

机构 * Aristotle University of Thessaloniki(亚里士多德大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12590 2025-09-19 cs.CV cs.LG 79%

Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering

Neale Ratzlaff, Matthew Lyle Olson, Musashi Hinck, Estelle Aflalo, Shao-Yen Tseng, Vasudev Lal, Phillip Howard

机构 * Oracle Intel Labs(英特尔实验室) Thoughtworks

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 6 Figures, 8 Tables. arXiv admin note: text overlap with arXiv:2410.13976

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14746 2025-09-19 cs.CV cs.IR 70%

Chain-of-Thought Re-ranking for Image Retrieval Tasks

Shangrong Wu, Yanghong Zhou, Yang Chen, Feng Zhang, P. Y. Mok

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 5 篇

2502.17651 2025-09-19 cs.CV cs.AI cs.CL 75%

METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling

Bingxuan Li, Yiwei Wang, Jiuxiang Gu, Kai-Wei Chang, Nanyun Peng

机构 * University of California, Los Angeles(加州大学洛杉矶分校) University of California, Merced(加州大学默塞德分校) Adobe Research(Adobe研究)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14638 2025-09-19 cs.CV 57%

MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks

Mingsong Li, Lin Liu, Hongjun Wang, Haoxing Chen, Xijun Gu, Shizhan Liu, Dong Gong, Junbo Zhao, Zhenzhong Lan, Jianguo Li

机构 * University of New South Wales(新南威尔士大学) The University of Hong Kong(香港大学) Zhejiang University(浙江大学) Westlake University(西湖大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14303 2025-09-19 cs.RO cs.AI 57%

FlowDrive: Energy Flow Field for End-to-End Autonomous Driving

Hao Jiang, Zhipeng Zhang, Yu Gao, Zhigang Sun, Yiru Wang, Yuwen Heng, Shuo Wang, Jinhao Chai, Zhuo Chen, Hao Zhao, Hao Sun, Xi Zhang, Anqing Jiang, Chuan Hu

机构 * Shanghai Jiao Tong University(上海交通大学) Bosch Corporate Research, Shanghai, China(博世股份有限公司上海研究院) AIR, Tsinghua University(清华大学人工智能研究院) Shanghai University, Shanghai, China(上海大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11277 2025-09-19 cs.CV cs.LG 57%

Probing the Representational Power of Sparse Autoencoders in Vision Models

Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff, Changbai Li, Phillip Howard, Vasudev Lal, Shao-Yen Tseng

机构 * Oracle Intel Labs(英特尔实验室) Oregon State University(俄勒冈州立大学) Thoughtworks

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14711 2025-09-19 eess.SP 50%

LLM4MG: Adapting Large Language Model for Multipath Generation via Synesthesia of Machines

Ziwei Huang, Shiliang Lu, Lu Bai, Xuesong Cai, Xiang Cheng

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 9 篇

2509.14886 2025-09-19 cs.CL cs.AI 84%

A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation

Ye Shen, Junying Wang, Farong Wen, Yijin Guo, Qi Jia, Zicheng Zhang, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学)

专题命中 多模态评测 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12772 2025-09-19 cs.CL cs.CV 84%

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, Ziwei Liu

机构 * LMMs-Lab Team(LMMs实验室团队) S-Lab, NTU, Singapore(新加坡国立大学S实验室)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Code ad leaderboard are available at https://github.com/EvolvingLMMs-Lab/lmms-eval and https://huggingface.co/spaces/lmms-lab/LiveBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15132 2025-09-19 cs.CY cs.CV 83%

From Pixels to Urban Policy-Intelligence: Recovering Legacy Effects of Redlining with a Multimodal LLM

Anthony Howell, Nancy Wu, Sharmistha Bagchi, Yushim Kim, Chayn Sun

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15222 2025-09-19 cs.SD cs.CV cs.MM eess.AS eess.IV 82%

Two Web Toolkits for Multimodal Piano Performance Dataset Acquisition and Fingering Annotation

Junhyung Park, Yonghyun Kim, Joonhyung Bae, Kirak Kim, Taegyun Kwon, Alexander Lerch, Juhan Nam

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to the Late-Breaking Demo Session of the 26th International Society for Music Information Retrieval (ISMIR) Conference, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏