arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-15 至 2025-09-15 共收录 27 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 1 篇

2504.04323 2025-09-15 cs.CV 57%

MedM-VL: What Makes a Good Medical LVLM?

Yiming Shi, Shaoshuai Yang, Xun Zhu, Haoyu Wang, Xiangling Fu, Miao Li, Ji Wu

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) College of AI, Tsinghua University(清华大学人工智能学院) Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心) School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2408.01284 2025-09-15 cs.MM cs.CV cs.SD eess.AS eess.IV 85%

Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework

Liuyuan Wen

机构 * School of Physical Sciences University of Science and Technology of China(中国科学技术大学物理科学学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to BMVC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09859 2025-09-15 cs.CV cs.LG 79%

WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector

Razvan Stefanescu, Ethan Oh, Ruben Vazquez, Chris Mesterharm, Constantin Serban, Ritu Chadha

机构 * Peraton Labs(珀顿实验室)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09747 2025-09-15 cs.LG cs.AI cs.RO 70%

D-CAT: Decoupled Cross-Attention Transfer between Sensor Modalities for Unimodal Inference

Leen Daher, Zhaobo Wang, Malcolm Mielle

机构 * Ecole Polytechnique Federale de Lausanne(瑞士联邦理工学院洛桑分校) Schindler EPFL Lab(Schindler EPFL实验室)

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 2 篇

2509.09804 2025-09-15 cs.CL 57%

Pragmatic Frames Evoked by Gestures: A FrameNet Brasil Approach to Multimodality in Turn Organization

Helen de Andrade Abreu, Tiago Timponi Torrent, Ely Edison da Silva Matos

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments Paper submitted to Language Sciences Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09889 2025-09-15 cs.RO cs.HC 50%

Using the Pepper Robot to Support Sign Language Communication

Giulia Botta, Marco Botta, Cristina Gena, Alessandro Mazzei, Massimo Donini, Alberto Lillo

专题命中 视频多模态 :multi-modal(abstract)

Comments paper presented at ICSR2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2509.09721 2025-09-15 cs.CV cs.AI cs.LG 84%

A Multimodal RAG Framework for Housing Damage Assessment: Collaborative Optimization of Image Encoding and Policy Vector Retrieval

Jiayi Miao, Dingxin Lu, Zhuqi Wang

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07879 2025-09-15 cs.IR cs.AI cs.CV 84%

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval

Wei Yang, Jingjing Fu, Rui Wang, Jinyu Wang, Lei Song, Jiang Bian

机构 * Microsoft Research Asia(微软亚洲研究院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10282 2025-09-15 cs.CV cs.LG 79%

MCL-AD: Multimodal Collaboration Learning for Zero-Shot 3D Anomaly Detection

Gang Li, Tianjiao Chen, Mingle Zhou, Min Li, Delong Han, Jin Wan

机构 * Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan)(计算能力网络与信息安全部教育部重点实验室,山东计算机科学中心(济南国家超级计算机中心)) Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东科学院)) Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science(山东省计算能力互联网与服务计算重点实验室,山东省计算机科学基础研究中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Page 14, 5 pictures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 1 篇

2509.09717 2025-09-15 cs.SD cs.LG eess.AS 57%

Testing chatbots on the creation of encoders for audio conditioned image generation

Jorge E. León, Miguel Carrasco

专题命中 多模态生成 :image-text(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 7 篇

2509.09730 2025-09-15 cs.CV cs.AI 84%

MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance

Kaikai Zhao, Zhaoxiang Liu, Peng Wang, Xin Wang, Zhicheng Ma, Yajun Xu, Wenjing Zhang, Yibing Nan, Kai Wang, Shiguo Lian

机构 * organization= Data Science \& Artificial Intelligence Research Institute, China Unicom , city= Beijing , postcode= 100033 , country= PR China

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments accepted by Image and Vision Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10059 2025-09-15 cs.CV cs.AI 81%

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration

Yue Zhou, Litong Feng, Mengcheng Lan, Xue Yang, Qingyun Li, Yiping Ke, Xue Jiang, Wayne Zhang

机构 * Department of Physics, J.K. Institute of Science(J.K.科学研究院物理系) World Scientific University(世界科学大学) University of Intelligent Studies(智能研究大学) East China Normal University(华东师范大学) Nanyang Technological University(南洋理工大学) SenseTime Research(商汤科技研究院) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 17 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09683 2025-09-15 cs.IR cs.AI 79%

Forecasting Clicks in Digital Advertising: Multimodal Inputs and Interpretable Outputs

Briti Gangopadhyay, Zhao Wang, Shingo Takamatsu

机构 * Sony Group Corporation(索尼集团公司) Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎- Rocquencourt 国家信息与自动化研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13818 2025-09-15 cs.CV cs.LG 79%

Building Age Estimation: A New Multi-Modal Benchmark Dataset and Community Challenge

Nikolaos Dionelis, Alessandra Feliciotti, Mattia Marconcini, Devis Peressutti, Nika Oman Kadunc, JaeWan Park, Hagai Raja Sinulingga, Steve Andreas Immanuel, Ba Tran, Caroline Arnold, Nicolas Longépé

机构 * European Space Agency(欧洲航天局) Φ \Phi -lab(Phi实验室) ESRIN(ESRIN研究所) MindEarth(MindEarth公司) Sinergise/ Planet(Sinergise/Planet公司) TelePIX(TelePIX公司) Axelspace Corporation(Axelspace公司) Helmholtz Institute Hereon(海德堡研究所) German Climate Computing Center DKRZ(德国气候计算中心DKRZ)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 20 figures, 1 table, Submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09722 2025-09-15 cs.CV cs.CL cs.LG 76%

Improving MLLM Historical Record Extraction with Test-Time Image

Taylor Archibald, Tony Martinez

机构 * Brigham Young University

专题命中 多模态评测 :MLLM(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11829 2025-09-15 cs.CL cs.AI 62%

Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation

Julia Kreutzer, Eleftheria Briakou, Sweta Agrawal, Marzieh Fadaee, Kocmi Tom

专题命中 多模态评测 :MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09799 2025-09-15 cs.LG cs.HC 50%

Distinguishing Startle from Surprise Events Based on Physiological Signals

Mansi Sharma, Alexandre Duchevet, Florian Daiber, Jean-Paul Imbert, Maurice Rekrut

机构 * German Research Center for Artificial Intelligence (DFKI), Cognitive Assistants Lab, Saarland Informatics Campus(德国人工智能研究中心(DFKI)、认知助理实验室、萨尔兰州信息技术校区) Fédération ENAC ISAE-SUPAERO ONERA, Université de Toulouse(ENAC ISAE-SUPAERO ONERA联合会、图卢兹大学)

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 2 篇

2508.03747 2025-09-15 cs.SI cs.AI cs.LG 57%

Data-Driven Discovery of Mobility Periodicity for Understanding Urban Systems

Xinyu Chen, Qi Wang, Yunhan Zheng, Nina Cao, HanQin Cai, Jinhua Zhao

机构 * Massachusetts Institute of Technology(麻省理工学院) Northeastern University(东北大学) University of Central Florida(佛罗里达中央大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02734 2025-09-15 eess.IV cs.CV cs.NE stat.AP stat.ML 57%

Integrative Variational Autoencoders for Generative Modeling of an Image Outcome with Multiple Input Images

Bowen Lei, Yeseul Jeon, Rajarshi Guhaniyogi, Aaron Scheffler, Bani Mallick, Alzheimer's Disease Neuroimaging Initiatives

机构 * Alzheimer’s Disease Neuroimaging Initiatives(阿尔茨海默病神经成像计划)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 多模态训练与对齐 6 篇

2411.02992 2025-09-15 cs.IR cs.CV 88%

Efficient and Effective Adaptation of Multimodal Foundation Models in Sequential Recommendation

Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Kaiwen Zheng, Yongxin Ni, Joemon M. Jose

机构 * School of Computing Science, University of Glasgow(格拉斯哥大学计算机科学学院) Amazon(亚马逊) Telefonica Scientific Research(Telefonica科学研究院) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Knowledge and Data Engineering (TKDE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09940 2025-09-15 cs.LG 88%

DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition

Yifei Wang, Wenbin Wang, Yong Luo

机构 * Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title);audio-visual(abstract)

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10408 2025-09-15 cs.CV cs.AI 81%

Multimodal SAM-adapter for Semantic Segmentation

Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli, Luigi Di Stefano

机构 * University of Bologna(博洛尼亚大学) SINA

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21135 2025-09-15 cs.CV cs.AI 81%

HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection

Harris Song, Tuan-Anh Vu, Sanjith Menon, Sriram Narasimhan, M. Khalid Jawed

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments fix typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17910 2025-09-15 cs.LG cs.AI 79%

A Novel Approach to Balance Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes and its Implementation in BEACON

Vansh Nagpal, Siva Likitha Valluru, Kausik Lakkaraju, Nitin Gupta, Zach Abdulrahman, Andrew Davison, Biplav Srivastava

机构 * BEACON

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10139 2025-09-15 cs.RO 71%

CaR1: A Multi-Modal Baseline for BEV Vehicle Segmentation via Camera-Radar Fusion

Santiago Montiel-Marín, Angel Llamazares, Miguel Antunes-García, Fabio Sánchez-García, Luis M. Bergasa

机构 * Department of Electronics. University of Alcalá, Community of Madrid, Spain(电子系。阿尔卡拉大学,马德里社区)

专题命中 多模态训练与对齐 :multi-modal(title)

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

9. 其他多模态 2 篇

2509.09805 2025-09-15 cs.RO 78%

MIMo grows! Simulating body and sensory development in a multimodal infant model

Francisco M. López, Miles Lenz, Marco G. Fedozzi, Arthur Aubret, Jochen Triesch

机构 * Deutsche Forschungsgemeinschaft(德国研究协会)

专题命中 其他多模态 :multimodal(title,abstract)

Comments Accepted at IEEE ICDL 2025. 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09869 2025-09-15 cs.CV cs.AI 62%

Surrogate Supervision for Robust and Generalizable Deformable Image Registration

Yihao Liu, Junyu Chen, Lianrui Zuo, Shuwen Wei, Brian D. Boyd, Carmen Andreescu, Olusola Ajilore, Warren D. Taylor, Aaron Carass, Bennett A. Landman

机构 * Department of Electrical and Computer Engineering, Vanderbilt University(维斯尼尔大学电气与计算机工程系) Department of Radiology and Radiological Science, Johns Hopkins Medical School(约翰霍普金斯医学学校放射学与放射科学系) Image Analysis and Communications Laboratory in the Department of Electrical and Computer Engineering, Johns Hopkins University(约翰霍普金斯大学电气与计算机工程系图像分析与通信实验室) Center for Cognitive Medicine, Department of Psychiatry and Behavioral Science, Vanderbilt University Medical Center(维斯尼尔大学医学中心认知医学中心) University of Pittsburgh, School of Medicine(匹兹堡大学医学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏