arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-26 至 2025-09-26 共收录 57 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2508.06434 2025-09-26 cs.CV cs.AI 86%

CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment

Shengzhu Yang, Jiawei Du, Shuai Lu, Weihang Zhang, Ningli Wang, Huiqi Li

机构 * Beijing Institute of Technology(北京理工大学) Beijing Tongren Hospital(北京同仁医院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07487 2025-09-26 cs.CV 85%

LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?

Bangyan Li, Wenxuan Huang, Zhenkun Gao, Yeqiang Wang, Yunhang Shen, Jingzhong Lin, Ling You, Yuxiang Shen, Shaohui Lin, Wanli Ouyang, Yuling Sun

机构 * East China Normal University(华东师范大学) The Chinese University of Hong Kong(香港中文大学) Northwest A&F University(西北农林科技大学) Tencent Youtu Lab(腾讯优图实验室) Xiamen University(厦门大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20961 2025-09-26 cs.CV cs.AI 84%

Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos

Sarmistha Das, R E Zera Marveen Lyngkhoi, Sriparna Saha, Alka Maurya

机构 * Indian Institute of Technology Patna(印度帕纳杰大学) CRISIL LTD(CRISIL公司)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20769 2025-09-26 cs.IR cs.AI cs.CV 81%

Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems

Tuo Zhang, Yuechun Sun, Ruiliang Liu

机构 * Museus University of Science and Technology of China(中国科学技术大学) British Museum(大英博物馆)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21251 2025-09-26 cs.CV cs.AI 76%

Instruction-tuned Self-Questioning Framework for Multimodal Reasoning

You-Won Jang, Yu-Jung Heo, Jaeseok Kim, Minsu Lee, Du-Seong Chang, Byoung-Tak Zhang

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments This paper was accepted to the "CLVL: 5th Workshop on Closing the Loop Between Vision and Language (ICCV 2023 CLVL workshop)."

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21287 2025-09-26 cs.CL cs.AI 73%

DisCoCLIP: A Distributional Compositional Tensor Network Encoder for Vision-Language Understanding

Kin Ian Lo, Hala Hawashin, Mina Abbaszadeh, Tilen Limback-Stokin, Hadi Wazni, Mehrnoosh Sadrzadeh

机构 * University College London(伦敦大学学院)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18174 2025-09-26 cs.CV cs.CL 73%

Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR

Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed, Ahmad Bastati, Zeina Aldallal, Sara Chrouf, Safwan AlModhayan

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00827 2025-09-26 cs.CV cs.AI 73%

IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves

Ruofan Wang, Juncheng Li, Yixu Wang, Bo Wang, Xiaosen Wang, Yan Teng, Yingchun Wang, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Huawei Technologies Ltd.(华为技术有限公司) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20792 2025-09-26 cs.CV cs.AI cs.LG 66%

DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation

Ved Umrajkar

机构 * Indian Institute of Technology, Roorkee(印度理工学院拉胡尔分校)

专题命中 图文多模态 :multimodal(abstract,comments);分类 cs.CV、cs.AI

Comments Accepted at ICCV2025 Workshop on Safe and Trustworthy Multimodal AI Systems

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2509.20724 2025-09-26 cs.SI cs.CL cs.CV cs.MM 82%

Visual Authority and the Rhetoric of Health Misinformation: A Multimodal Analysis of Social Media Videos

Mohammad Reza Zarei, Barbara Stead-Coyle, Michael Christensen, Sarah Everts, Majid Komeili

机构 * School of Computer Science(计算机科学学院) Carleton University(卡尔顿大学) Department of Law and Legal Studies(法律与法律研究系) School of Journalism and Communication(新闻与传播学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20741 2025-09-26 eess.AS cs.ET cs.LG 79%

Real-Time System for Audio-Visual Target Speech Enhancement

T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang

机构 * Bose Corporation(博世公司) Georgia Institute of Technology(佐治亚理工学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted into WASPAA 2025 demo session

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20467 2025-09-26 cs.CL cs.CV 62%

ShortCheck: Checkworthiness Detection of Multilingual Short-Form Videos

Henrik Vatndal, Vinay Setty

机构 * Factiverse AI University of Stavanger(斯塔万格大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21144 2025-09-26 cs.SD cs.AI 57%

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

Sitong Cheng, Weizhen Bian, Xinsheng Wang, Ruibin Yuan, Jianyi Chen, Shunshun Yin, Yike Guo, Wei Xue

机构 * Hong Kong University of Science and Technology(香港理工大学) Soul AI Lab(Soul AI 实验室)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07282 2025-09-26 eess.AS 57%

Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild

Jing-Tong Tzeng, Bo-Hao Su, Ya-Tse Wu, Hsing-Hang Chou, Chi-Chun Lee

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Proceedings of Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 8 篇

2509.21100 2025-09-26 cs.CV 83%

VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception

Ziang Yan, Xinhao Li, Yinan He, Zhengrong Yue, Xiangyu Zeng, Yali Wang, Yu Qiao, Limin Wang, Yi Wang

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室) Nanjing University(南京大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19973 2025-09-26 cs.CV 79%

OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving

Pei Liu, Hongliang Lu, Haichao Liu, Haipeng Liu, Xin Liu, Ruoyu Yao, Shengbo Eben Li, Jun Ma

机构 * The Hong Kong University of Science and Technology(香港科技大学) Li Auto Inc. the School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02781 2025-09-26 q-bio.QM cs.AI cs.LG 79%

Multimodal AI predicts clinical outcomes of drug combinations from preclinical data

Yepeng Huang, Xiaorui Su, Varun Ullanat, Intae Moon, Ivy Liang, Lindsay Clegg, Damilola Olabode, Ruthie Johnson, Nicholas Ho, Megan Gibbs, Megan Gibbs, Alexander Gusev, Bino John, Marinka Zitnik

机构 * Harvard Medical School(哈佛医学院) Harvard College(哈佛学院) AstraZeneca(阿斯利康) Carnegie Mellon University(卡内基梅隆大学) Dana-Farber Cancer Institute and Harvard Medical School(达纳-法伯癌症研究所和哈佛医学院) Harvard University(哈佛大学) Broad Institute of MIT and Harvard(MIT和哈佛大学 Broad研究所) Harvard Data Science Initiative(哈佛数据科学计划)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20703 2025-09-26 cs.RO cs.AI cs.CV 62%

Joint Flow Trajectory Optimization For Feasible Robot Motion Generation from Video Demonstrations

Xiaoxiang Dong, Matthew Johnson-Roberson, Weiming Zhi

机构 * College of Connected Computing, Vanderbilt University(连接计算学院,范德比尔特大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学) School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16421 2025-09-26 cs.CV cs.AI 62%

AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead

Aiden Chang, Celso De Melo, Stephanie M. Lukin

机构 * University of Southern California(南加州大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at NeurIPS 2025, 32 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17844 2025-09-26 cs.CL cs.AI 62%

THCM-CAL: Temporal-Hierarchical Causal Modelling with Conformal Calibration for Clinical Risk Prediction

Xin Zhang, Qiyu Wei, Yingjie Zhu, Fanyi Wu, Sophia Ananiadou

机构 * The University of Manchester(曼彻斯特大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18056 2025-09-26 cs.CV 57%

TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs

Yunheng Li, Jing Cheng, Shaoyong Jia, Hangyi Kuang, Shaohui Jiao, Qibin Hou, Ming-Ming Cheng

机构 * VCIP, School of Computer Science, Nankai University(VCIP,计算机科学学院,南开大学) ByteDance Inc.(字节跳动公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21851 2025-09-26 cs.RO cs.AI cs.LG 57%

Streaming Flow Policy: Simplifying diffusion/flow-matching policies by treating action trajectories as flow trajectories

Sunshine Jiang, Xiaolin Fang, Nicholas Roy, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Siddharth Ancha

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Conference on Robot Learning (CoRL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 6 篇

2509.03477 2025-09-26 cs.LG cs.AI cs.CV 81%

Robult: Leveraging Redundancy and Modality Specific Features for Robust Multimodal Learning

Duy A. Nguyen, Abhi Kamboj, Minh N. Do

机构 * Siebel School of Computing and Data Science, UIUC, US(UIUC计算机与数据科学学院) Department of Electrical and Computer Engineering, UIUC, US(UIUC电气与计算机工程学院) VinUni-Illinois Smart Health Center, VinUniversity, Vietnam(Vin大学-伊利诺伊智能健康中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted and presented at IJCAI 2025 in Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21151 2025-09-26 cs.CL cs.IR 79%

Retrieval over Classification: Integrating Relation Semantics for Multimodal Relation Extraction

Lei Hei, Tingjing Liao, Yingxin Pei, Yiyang Qi, Jiaqi Wang, Ruiting Li, Feiliang Ren

机构 * School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China(计算机科学与工程学院,东北大学,沈阳110819,中国)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20501 2025-09-26 cs.LG cs.CV 79%

Beyond Visual Similarity: Rule-Guided Multimodal Clustering with explicit domain rules

Kishor Datta Gupta, Mohd Ariful Haque, Marufa Kamal, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Roy George

机构 * Clark Atlanta University(克拉克阿特兰大学) BRAC University(布拉克大学) United International University(国际联合大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15141 2025-09-26 cs.AI cs.LG physics.chem-ph 79%

Text-Augmented Multimodal LLMs for Chemical Reaction Condition Recommendation

Yu Zhang, Ruijie Yu, Kaipeng Zeng, Ding Li, Feng Zhu, Xiaokang Yang, Yaohui Jin, Yanyan Xu

机构 * Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎-罗克琴克研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20813 2025-09-26 cs.CV cs.AI 73%

Revolutionizing Precise Low Back Pain Diagnosis via Contrastive Learning

Thanh Binh Le, Hoang Nhat Khang Vo, Tan-Ha Mai, Trong Nhan Phan

机构 * Faculty of Computer Science and Engineering(计算机科学与工程学院) Ho Chi Minh City University of Technology(胡志明市技术大学) Vietnam National University (HCMUT-VNU)(越南国家大学(HCMUT-VNU)) National Taiwan University(国立台湾大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21259 2025-09-26 cs.NI cs.AI 57%

Semantic Edge-Cloud Communication for Real-Time Urban Traffic Surveillance with ViT and LLMs over Mobile Networks

Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 9 篇

2505.15809 2025-09-26 cs.CV 83%

MMaDA: Multimodal Large Diffusion Language Models

Ling Yang, Ye Tian, Bowen Li, Xinchen Zhang, Ke Shen, Yunhai Tong, Mengdi Wang

机构 * Princeton University(普林斯顿大学) Peking University(北京大学) Tsinghua University(清华大学) ByteDance Seed(字节跳动种子)

专题命中 多模态生成 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments NeurIPS 2025. Project: https://github.com/Gen-Verse/MMaDA

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12148 2025-09-26 cs.CV 83%

HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation

Ling Yang, Xinchen Zhang, Ye Tian, Chenming Shang, Minghao Xu, Wentao Zhang, Bin Cui

机构 * Peking University(北京大学) Tsinghua University(清华大学) Mila - Québec AI Institute(魁北克人工智能研究所)

专题命中 多模态生成 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments NeurIPS 2025. Code: https://github.com/Gen-Verse/HermesFlow

详情

展开后加载摘要…

URL PDF HTML 收藏