arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-23 至 2025-10-23 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 5 篇

2510.19559 2025-10-23 cs.CV cs.AI cs.IR cs.MM 67%

A Matter of Time: Revealing the Structure of Time in Vision-Language Models

Nidham Tekaya, Manuela Waldner, Matthias Zeppelzauer

机构 * St. Pölten University of Applied Sciences(施普伦特应用科学大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19678 2025-10-23 cs.CV cs.AI 62%

I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs

John Burden, Jonathan Prunty, Ben Slater, Matthieu Tehenan, Greg Davis, Lucy Cheke

机构 * Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能研究中心、剑桥大学) Department of Engineering, University of Cambridge(工程系、剑桥大学) Department of Psychology, University of Cambridge(心理学系、剑桥大学) Department of Computer Science, University of Cambridge(计算机科学系、剑桥大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19001 2025-10-23 cs.CV cs.AI cs.RO 62%

Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts

Seungjun Yu, Junsung Park, Youngsun Lim, Hyunjung Shim

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19802 2025-10-23 cs.CV 57%

Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models

Xiaozhen Qiao, Jingkai Zhao, Yuqiu Jiang, Xianda Guo, Zhe Sun, Hongyuan Zhang, Xuelong Li

机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Institute of Artificial Intelligence (TeleAI), China Telecom, P. R. China(人工智能研究所(TeleAI),中国电信,中华人民共和国) College of Computer Science, Wuhan University(计算机科学学院,武汉大学) School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院,光学与电子学(iOPEN),西北工业大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16814 2025-10-23 cs.LG cs.CV 57%

Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning

Junhao Shen, Haiteng Zhao, Yuzhe Gu, Songyang Gao, Kuikun Liu, Haian Huang, Jianfei Gao, Dahua Lin, Wenwei Zhang, Kai Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2510.19358 2025-10-23 cs.CL cs.AI 84%

M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models

Yejin Kwon, Taewoo Kang, Hyunsoo Yoon, Changouk Kim

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments Submitted to LREC 2026. 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00937 2025-10-23 cs.DC cs.AI 79%

ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving

Haoran Qiu, Anish Biswas, Zihan Zhao, Jayashree Mohan, Alind Khare, Esha Choukse, Íñigo Goiri, Zeyu Zhang, Haiying Shen, Chetan Bansal, Ramachandran Ramjee, Rodrigo Fonseca

机构 * Microsoft Azure Research(微软Azure研究) Microsoft Research India(微软印度研究) University of Virginia(弗吉尼亚大学) Microsoft M365 Research(微软M365研究)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published at ACM SoCC 2025; 14 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19055 2025-10-23 cs.AI cs.SD eess.AS 62%

The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS

Brandon James Carone, Iran R. Roman, Pablo Ripollés

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 5 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23670 2025-10-23 cs.SD cs.CL eess.AS 62%

Efficient Interleaved Speech Modeling through Knowledge Distillation

Mohammadmahdi Nouriborji, Morteza Rohanian

机构 * Nlpie Research University of Zurich(Nlpie研究所 瑞士苏黎世大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19144 2025-10-23 cs.CL 57%

Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges

Cheng Huang, Nyima Tashi, Fan Gao, Yutong Liu, Jiahao Li, Hao Tian, Siyang Jiang, Thupten Tsering, Ban Ma-bao, Renzeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Jin Zhang, Xiao Feng, Hao Wang, Jie Tang, Guojie Tang, Xiangxiang Wang, Jia Zhang, Tsengdar Lee, Yongbin Yu

机构 * University of Electronic Science and Technology of China(电子科技大学) Southern Methodist University(南方 Methodist 大学) The City University of Hong Kong(香港城市大学) The Hong Kong Polytechnic University(香港理工大学) The Chinese University of Hong Kong(香港中文大学) University of Connecticut(康涅狄格大学) Tsinghua University(清华大学) University of Texas at Arlington(德克萨斯大学阿灵顿分校) University of Chinese Academy of Sciences(中国科学院大学) Tibet University(西藏大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2505.24625 2025-10-23 cs.CV cs.AI 73%

Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors

Duo Zheng, Shijia Huang, Yanyang Li, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16708 2025-10-23 cs.CL cs.AI 62%

Natural Language Processing for Cardiology: A Narrative Review

Kailai Yang, Yan Leng, Xin Zhang, Tianlin Zhang, Paul Thompson, Bernard Keavney, Maciej Tomaszewski, Sophia Ananiadou

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19574 2025-10-23 cs.CV cs.CR 57%

Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection

Ariana Yi, Ce Zhou, Liyang Xiao, Qiben Yan

机构 * Mission San Jose High School(Mission San Jose 高中) Missouri University of Science and Technology(密苏里科学与技术大学) Michigan State University(密歇根州立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19560 2025-10-23 cs.CV 57%

HAD: Hierarchical Asymmetric Distillation to Bridge Spatio-Temporal Gaps in Event-Based Object Tracking

Yao Deng, Xian Zhong, Wenxuan Liu, Zhaofei Yu, Jingling Yuan, Tiejun Huang

机构 * Sanya Science and Education Innovation Park, Wuhan University of Technology(武汉理工大学三亚科学教育创新园) Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology(湖北省交通运输物联网重点实验室,计算机科学与人工智能学院,武汉理工大学) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21776 2025-10-23 cs.CV 57%

Video-R1: Reinforcing Video Reasoning in MLLMs

Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, Xiangyu Yue

机构 * CUHK MMLab(香港中文大学多模态实验室) CUHK (SZ)(香港中文大学(深圳)) Tsinghua University(清华大学) UCAS(中国科学院大学) CUHK HCCL(香港中文大学高性能计算实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025, Project page: https://github.com/tulerfeng/Video-R1

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2510.19398 2025-10-23 cs.CL 57%

SONAR-SLT: Multilingual Sign Language Translation via Language-Agnostic Sentence Embedding Supervision

Yasser Hamidullah, Shakib Yazdani, Cennet Oguz, Josef van Genabith, Cristina España-Bonet

机构 * German Research Center for Artificial Intelligence (DFKI GmbH)(德国人工智能研究中心(DFKI GmbH)) Saarland Informatics Campus(萨尔兰信息技术校区) Barcelona Supercomputing Center (BSC-CNS)(巴塞罗那超级计算中心(BSC-CNS))

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Journal ref published at WMT2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 9 篇

2510.18879 2025-10-23 cs.HC 78%

FIRETWIN: Digital Twin Advancing Multi-Modal Sensing, Interactive Analytics for Wildfire Response

Mayamin Hamid Raha, Ali Reza Tavakkoli, Chris Webb, Mobin Habibpour, Janice Coen, Eric Rowell, Fatemeh Afghah

专题命中 多模态生成 :multi-modal(title);multimodal(abstract)

Comments 8 pages, 6 figures, accepted in IEEE International Workshop on Computer-Aided Modeling and Design of Communication Links and Networks (CAMAD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19808 2025-10-23 cs.CV cs.CL cs.LG 73%

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

Yusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong, Yinfei Yang, Jiasen Lu, Wenze Hu, Zhe Gan

机构 * Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19641 2025-10-23 cs.CL cs.AI 62%

Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent

Yangshijie Zhang, Xinda Wang, Jialin Liu, Wenqiang Wang, Zhicong Ma, Xingxing Jia

机构 * Lanzhou University(兰州大学) Peking University(北京大学) Sun Yat-sen University(中山大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17519 2025-10-23 cs.CV cs.AI 62%

MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models

Yongshun Zhang, Zhongyi Fan, Yonghang Zhang, Zhangzikang Li, Weifeng Chen, Zhongwei Feng, Chaoyue Wang, Peng Hou, Anxiang Zeng

机构 * LLM Team, Shopee Pte. Ltd.(Shopee 股份有限公司语言模型团队)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Technical Report; Project Page: https://github.com/Shopee-MUG/MUG-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00939 2025-10-23 cs.CV cs.CL 62%

WikiVideo: Article Generation from Multiple Videos

Alexander Martin, Reno Kriz, William Gantt Walden, Kate Sanders, Hannah Recknor, Eugene Yang, Francis Ferraro, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学) Human Language Technology Center of Excellence(人机语言技术卓越中心) University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Repo can be found here: https://github.com/alexmartin1722/wikivideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18911 2025-10-23 physics.chem-ph cs.AI 57%

Prospects for Using Artificial Intelligence to Understand Intrinsic Kinetics of Heterogeneous Catalytic Reactions

Andrew J. Medford, Todd N. Whittaker, Bjarne Kreitz, David W. Flaherty, John R. Kitchin

机构 * organization= School of Chemical \& Biomolecular Engineering, Georgia Institute of Technology , addressline= 311 Ferst Drive NW , city= Atlanta , postcode= 30332 , state= GA , country= USA organization= Department of Chemical Engineering, Carnegie Mellon University , addressline= 5000 Forbes Street , city= Pittsburgh , postcode= 15213 , state= PA , country= USA

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Submitted to "Current Opinion in Chemical Engineering" for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12323 2025-10-23 cs.CV 57%

Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert Feedback

Janet Wang, Yunbei Zhang, Zhengming Ding, Jihun Hamm

机构 * Tulane University(路易斯安那州立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05051 2025-10-23 cs.CV cs.RO 57%

ComDrive: Comfort-Oriented End-to-End Autonomous Driving

Junming Wang, Xingyu Zhang, Zebin Xing, Songen Gu, Xiaoyang Guo, Yang Hu, Ziying Song, Qian Zhang, Xiaoxiao Long, Wei Yin

机构 * Horizon Robotics University of Hong Kong(香港大学) University of the Chinese Academy of Sciences(中国科学院大学) Nanjing University(南京大学) Beijing Jiaotong University(北京交通大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12407 2025-10-23 cs.DC cs.LG 50%

The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution

Frank Sifei Luan, Ron Yifeng Wang, Yile Gu, Ziming Mao, Charlotte Lin, Amog Kamsetty, Hao Chen, Cheng Su, Balaji Veeramani, Scott Lee, SangBin Cho, Clark Zinzow, Eric Liang, Ion Stoica, Stephanie Wang

机构 * UC Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) Anyscale

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 11 篇

2506.16962 2025-10-23 cs.CV cs.AI cs.CL 85%

Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search

Haoran Sun, Yankai Jiang, Wenjie Lou, Yujie Zhang, Wenjie Li, Lilong Wang, Mianxin Liu, Lei Liu, Xiaosong Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19451 2025-10-23 cs.CV cs.MM 81%

Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis

Xueqi Ma, Yanbei Jiang, Sarah Erfani, James Bailey, Weifeng Liu, Krista A. Ehinger, Jey Han Lau

机构 * The University of Melbourne(墨尔本大学) China University of Petroleum (East China)(中国石油大学(华东))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19183 2025-10-23 cs.CV cs.AI 81%

PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning

Fengyuan Sun, Hui Chen, Xinhao Xu, Dandan Zheng, Jingdong Chen, Jun Zhou, Jungong Han, Guiguang Ding

机构 * School of Software, Tsinghua University(清华大学软件学院) Ant Group(蚂蚁集团) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00711 2025-10-23 cs.LG cs.AI cs.CV 81%

QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training

Wei Dai, Peilin Chen, Chanakya Ekbote, Paul Pu Liang

机构 * MIT Media Lab(MIT媒体实验室) MIT EECS(MIT电子工程与计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted as Oral at NeurIPS 2025. Revision after camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19305 2025-10-23 cs.LG cs.CV 79%

FrogDeepSDM: Improving Frog Counting and Occurrence Prediction Using Multimodal Data and Pseudo-Absence Imputation

Chirag Padubidri, Pranesh Velmurugan, Andreas Lanitis, Andreas Kamilaris

机构 * Pervasive Systems, University of Twente(普罗威斯系统,埃因霍温大学) CYENS Center of Excellence(CYENS卓越中心) Cyprus University of Tecnhology(塞浦路斯技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏