arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-01 至 2025-10-01 共收录 77 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2509.25717 2025-10-01 cs.CV cs.CL cs.LG 81%

Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization

Xintong Li, Chuhan Wang, Junda Wu, Rohan Surana, Tong Yu, Julian McAuley, Jingbo Shang

机构 * University of California, San Diego(加州大学圣迭戈分校) Adobe Research(Adobe研究)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22820 2025-10-01 cs.CV cs.AI 81%

MMPB: It's Time for Multi-Modal Personalization

Jaeik Kim, Woojin Kim, Woohyeon Park, Jaeyoung Do

机构 * AIDAS Laboratory(AIDAS实验室) IPAI ECE(电子工程系) Seoul National University(首尔国立大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25791 2025-10-01 cs.CV 79%

EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks

Yuan Gao, Sangwook Kim, Chris McIntosh

机构 * Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络) Department of Medical Biophysics, UofT(医学生物物理学系) Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络) Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学) Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute) Department of Medical Imaging, UofT(医学影像学系) Vector Institute, Toronto(向量研究所)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments MICCAI 2025

Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. MICCAI 2025. Lecture Notes in Computer Science, vol 15964. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19294 2025-10-01 cs.CV cs.AI cs.CL 78%

Object Detection with Multimodal Large Vision-Language Models: An In-depth Review

Ranjan Sapkota, Manoj Karkee

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments First Peer Reviewed Review Paper for Object Detection with Vision-Language Models (VLMs)

Journal ref Information Fusion, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25654 2025-10-01 cs.CV 70%

DescribeEarth: Describe Anything for Remote Sensing Images

Kaiyu Li, Zixuan Jiang, Xiangyong Cao, Jiayu Wang, Yuchen Xiao, Deyu Meng, Zhi Wang

机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室) School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室) Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)

专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06105 2025-10-01 cs.CV 70%

PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology

Yating Huang, Ziyan Huang, Lintao Xiang, Qijun Yang, Hujun Yin

机构 * University of Manchester(曼彻斯特大学) South China University of Technology(华南理工大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accept by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25202 2025-10-01 cs.LG 67%

VLHSA: Vision-Language Hierarchical Semantic Alignment for Jigsaw Puzzle Solving with Eroded Gaps

Zhuoning Xu, Xinyan Liu

机构 * Zhuoning Xu 1(Xu Zhuoning 1) Xinyan Liu 1(Liu Xinyan 1)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07675 2025-10-01 cs.LG cs.AI cs.CV 62%

Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization

Seongjae Kang, Dong Bok Lee, Hyungjoon Jang, Sung Ju Hwang

机构 * VUNO Inc.(VUNO公司) KAIST(韩国科学技术院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 38 pages, 17 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21432 2025-10-01 cs.RO cs.CV 57%

UAV-VLN: End-to-End Vision Language guided Navigation for UAVs

Pranav Saxena, Nishant Raghuvanshi, Neena Goveas

机构 * Birla Institute of Technology and Science Pilani, K.K Birla Goa Campus(比拉理工学院和科学学院,K.K比拉果阿校区)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Journal ref Proc. European Conference on Mobile Robots (ECMR), 2025, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 9 篇

2505.20873 2025-10-01 cs.CV 88%

Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models

Chaeyoung Jung, Youngjoon Jang, Jongmin Choi, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20862 2025-10-01 cs.CV 85%

AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

Chaeyoung Jung, Youngjoon Jang, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25652 2025-10-01 cs.AI cs.MM cs.SD 84%

Iterative Residual Cross-Attention Mechanism: An Integrated Approach for Audio-Visual Navigation Tasks

Hailong Zhang, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng

机构 * Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center, School of Computer Science and Technology, Xinjiang University(新疆多模态智能处理与信息安全工程技术创新中心,计算机科学与技术学院,新疆大学) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) School of Electrical Engineering and Automation, Tianjin University of Technology(电气工程与自动化学院,天津工业大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM

Comments Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15772 2025-10-01 cs.SD cs.CL eess.AS 84%

MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling

Yifan Cheng, Ruoyi Zhang, Jiatong Shi

机构 * Santa Clara, CA, USA(美国圣克拉拉) Carnegie Mellon University(卡内基梅隆大学) Huazhong University of Science and Technology(华中科技大学) Nanjing University of Information Science and Technology(南京信息工程大学)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CL、eess.AS

Comments Accepted by Interspeech

Journal ref Proc. of Interspeech2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13624 2025-10-01 cs.SD eess.AS 79%

Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement

Rong Chao, Wenze Ren, You-Jin Li, Kuo-Hsuan Hung, Sung-Feng Huang, Szu-Wei Fu, Wen-Huang Cheng, Yu Tsao

机构 * Academia Sinica(台湾“中央研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted to Interspeech 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26604 2025-10-01 cs.CV 57%

Video Object Segmentation-Aware Audio Generation

Ilpo Viertola, Vladimir Iashin, Esa Rahtu

机构 * Tampere University, Tampere, Finland(塔尔库大学) University of Oxford, Oxford, UK(牛津大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Preprint version. The Version of Record is published in DAGM GCPR 2025 proceedings with Springer Lecture Notes in Computer Science (LNCS). Updated results and resources are available at the project page: https://saganet.notion.site

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25670 2025-10-01 cs.SD cs.CV 57%

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

Kang Yang, Yifan Liang, Fangkun Liu, Zhenping Xie, Chengshi Zheng

机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) Institute of Acoustics Chinese Academy of Science(中国科学院声学研究所)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26593 2025-10-01 cs.HC 50%

Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation

Kichang Lee

专题命中 音频语音多模态 :multimodal(abstract)

Comments 23 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14608 2025-10-01 cs.RO 50%

Visual-auditory Extrinsic Contact Estimation

Xili Yi, Jayjun Lee, Nima Fazeli

机构 * Robotics Department, University of Michigan(密歇根大学机器人系)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 10 篇

2509.26473 2025-10-01 cs.AI 83%

STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models

Shaoxiong Guo, Tianyi Du, Lijun Li, Yuyao Wu, Jie Li, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) East China Normal University(华东师范大学) Soochow University(苏州大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07032 2025-10-01 cs.CL cs.CV 81%

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin'ichi Satoh, Michael Felsberg, Mubarak Shah, Salman Khan, Fahad Shahbaz Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫德赫·本·扎耶德人工智能大学) University of Central Florida(中央佛罗里达大学) Islamic University of Technology(伊斯兰技术大学) Air University(空军大学) ETH Zurich(苏黎世联邦理工学院) Technische Universität München(慕尼黑技术大学) National Institute of Informatics(国家信息研究所) Australian National University(澳大利亚国立大学) Linköping University(利尔贝里大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26636 2025-10-01 cs.LG 78%

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

Shangding Gu, Xiaohan Wang, Donghao Ying, Haoyu Zhao, Runing Yang, Ming Jin, Boyi Li, Marco Pavone, Serena Yeung-Levy, Jun Wang, Dawn Song, Costas Spanos

机构 * UC Berkeley(伯克利大学) Stanford(斯坦福大学) UCL(伦敦大学学院) Virginia Tech(弗吉尼亚理工学院) Nvidia(英伟达公司)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00139 2025-10-01 cs.CV 74%

SuperEvent: Cross-Modal Learning of Event-based Keypoint Detection for SLAM

Yannick Burkhardt, Simon Schaefer, Stefan Leutenegger

机构 * Technical University of Munich(慕尼黑技术大学) ETH Zürich(苏黎世联邦理工学院) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24200 2025-10-01 cs.CV 70%

UniVid: The Open-Source Unified Video Model

Jiabin Luo, Junhui Lin, Zeyu Zhang, Biao Wu, Meng Fang, Ling Chen, Hao Tang

机构 * Peking University(北京大学) AI Geeks Australian Artificial Intelligence Institute(澳大利亚人工智能研究所)

专题命中 视频多模态 :MLLM(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25745 2025-10-01 cs.CV cs.CL cs.MM 67%

FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos

Siddhant Sukhani, Yash Bhardwaj, Riya Bhadani, Veer Kejriwal, Michael Galarnyk, Sudheer Chava

机构 * Stanford University(斯坦福大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments ICCV Short Video Understanding Workshop Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07966 2025-10-01 cs.CV cs.AI cs.CL 67%

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu, Hanrong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, Sifei Liu, Hongxu Yin, Yao Lu, Song Han

机构 * NVIDIA MIT(麻省理工学院) HKU(香港大学) UC Berkeley(加州大学伯克利分校)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS 2025. Code at https://github.com/NVlabs/Long-RL and model at https://huggingface.co/Efficient-Large-Model/LongVILA-R1-7B

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26225 2025-10-01 cs.CV cs.AI 62%

An Experimental Study on Generating Plausible Textual Explanations for Video Summarization

Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris

机构 * IEEE CBMI 2025

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments IEEE CBMI 2025. This is the authors' accepted version. The final publication is available at https://ieeexplore.ieee.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04909 2025-10-01 cs.CV cs.AI 62%

HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding

Yuxuan Cai, Jiangning Zhang, Zhenye Gan, Qingdong He, Xiaobin Hu, Junwei Zhu, Yabiao Wang, Chengjie Wang, Zhucun Xue, Chaoyou Fu, Xinwei He, Xiang Bai

机构 * Fantasyele

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01894 2025-10-01 cs.CV cs.HC 57%

Adaptive Modality Balanced Online Knowledge Distillation for Brain-Eye-Computer based Dim Object Detection

Zixing Li, Chao Yan, Zhen Lan, Xiaojia Xiang, Han Zhou, Jun Lai, Dengqing Tang

机构 * College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 18 pages,15 figures

Journal ref IEEE Transactions on Neural Networks and Learning Systems, pp. 1-15, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2509.25638 2025-10-01 cs.CV cs.LG 85%

Generalized Contrastive Learning for Universal Multimodal Retrieval

Jungsoo Lee, Janghoon Cho, Hyojin Park, Munawar Hayat, Kyuwoong Hwang, Fatih Porikli, Sungha Choi

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25711 2025-10-01 cs.CV 79%

ProbMed: A Probabilistic Framework for Medical Multimodal Binding

Yuan Gao, Sangwook Kim, Jianzhong You, Chris McIntosh

机构 * Peter Munk Cardiac Centre(彼得·默克心脏中心) Ted Rogers Centre for Heart Research(泰德·罗杰斯心脏病研究中心) University Health Network(大学健康网络) Joint Department of Medical Imaging(联合医学影像部门) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏