arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-10 至 2025-11-10 共收录 38 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2504.17902 2025-11-10 cs.CV cs.CL 81%

TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection

Girish A. Koushik, Helen Treharne, Aditya Joshi, Diptesh Kanojia

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to Special Track on AI for Social Impact (AISI) at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15418 2025-11-10 cs.CL cs.AI 81%

Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs

Lee Qi Zun, Mohamad Zulhilmi Bin Abdul Halim, Goh Man Fye

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11350 2025-11-10 cs.RO 78%

Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild

Derek Ming Siang Tan, Shailesh, Boyang Liu, Alok Raj, Qi Xuan Ang, Weiheng Dai, Tanishq Duhan, Jimmy Chiun, Yuhong Cao, Florian Shkurti, Guillaume Sartoretti

专题命中 图文多模态 :multimodal(title,abstract)

Comments Accepted for presentation at CORL 2025. Code, models, and data are available at https://search-tta.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05057 2025-11-10 cs.CV 70%

Role-SynthCLIP: A Role Play Driven Diverse Synthetic Data Approach

Yuanxiang Huangfu, Chaochao Wang, Weilei Wang

机构 * PatSnap Co., LTD.(PatSnap公司)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02615 2025-11-10 cs.LG 67%

ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models

Duy M. H. Nguyen, Nghiem T. Diep, Trung Q. Nguyen, Hoang-Bao Le, Tai Nguyen, Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, James Zou, Daniel Sonntag, Mathias Niepert

机构 * German Research Centre for Artificial Intelligence (DFKI)(德国人工智能研究中心) Max Planck Research School for Intelligent Systems (IMPRS-IS)(马克斯·普朗克智能系统研究学校) University of Stuttgart(斯图加特大学) University Medical Center Gottingen(哥廷根大学医学中心) Max Planck Institute for Multidisciplinary Sciences(马克斯·普朗克多学科科学研究所) ARC Centre of Excellence for the Mathematical Analysis of Cellular Systems(细胞系统数学分析卓越中心) School of Mathematical Sciences, Queensland University of Technology(昆士兰科技大学数学科学学院) University of Oldenburg(奥尔登堡大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of California San Diego(加州大学圣地亚哥分校) MBZUAI(马克斯·普朗克人工智能研究所) ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学)

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract)

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04247 2025-11-10 cs.MM cs.AI cs.IR 62%

On the Brittleness of CLIP Text Encoders

Allie Tran, Luca Rossetto

机构 * Dublin City University(都柏林城市大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI、cs.MM

Comments Accepted for publication at MMM'26. Analysis code can be found here: https://github.com/allie-tran/clip-brittleness

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2511.05432 2025-11-10 cs.CV 79%

Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis

Dogucan Yaman, Seymanur Akti, Fevziye Irem Eyiokur, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05304 2025-11-10 cs.HC 78%

psiUnity: A Platform for Multimodal Data-Driven XR

Akhil Ajikumar, Sahil Mayenkar, Steven Yoo, Sakib Reza, Mohsen Moghaddam

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04995 2025-11-10 cs.HC cs.AI cs.CL 62%

Enhancing Public Speaking Skills in Engineering Students Through AI

Amol Harsh, Brainerd Prince, Siddharth Siddharth, Deepan Raj Prabakar Muthirayan, Kabir S Bhalla, Esraaj Sarkar Gupta, Siddharth Sahu

机构 * Center for Thinking, Language and Communication(思考、语言与交流中心) Plaksha University(普拉克斯大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 3 篇

2411.00696 2025-11-10 cs.LG cs.AI 88%

CTPD: Cross-Modal Temporal Pattern Discovery for Enhanced Multimodal Electronic Health Records Analysis

Fuying Wang, Feng Wu, Yihan Tang, Lequan Yu

机构 * Department of Statistics and Actuarial Science, School of Computing and Data Science, The University of Hong Kong(统计与精算学系,计算与数据科学学院,香港大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09048 2025-11-10 cs.LG 71%

Spatio-Temporal Graph Convolutional Networks for EV Charging Demand Forecasting Using Real-World Multi-Modal Data Integration

Jose Tupayachi, Mustafa C. Camur, Kevin Heaslip, Xueping Li

机构 * University of Tennessee(田纳西大学)

专题命中 视频多模态 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05636 2025-11-10 cs.LG cs.AI 57%

Graph Learning

Feng Xia, Ciyuan Peng, Jing Ren, Falih Gozi Febrinanto, Renqiang Luo, Vidya Saikrishna, Shuo Yu, Xiangjie Kong

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 185 pages

Journal ref Foundations and Trends in Signal Processing, Vol. 19, No. 4, pp 371-551. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 7 篇

2505.16470 2025-11-10 cs.IR cs.CL cs.CV 84%

Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering

Kuicai Dong, Yujing Chang, Shijie Huang, Yasheng Wang, Ruiming Tang, Yong Liu

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Paper accepted to NeurIPS 2025 DB

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05325 2025-11-10 cs.LG 82%

Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval

Janet Jenq, Hongda Shen

机构 * PitchBook USA(PitchBook美国公司)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08828 2025-11-10 cs.IR cs.AI cs.CL cs.CV 82%

MMDocIR: Benchmarking Multimodal Retrieval for Long Documents

Kuicai Dong, Yujing Chang, Xin Deik Goh, Dexun Li, Ruiming Tang, Yong Liu

机构 * Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Paper accepted to EMNLP-2025(Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05404 2025-11-10 cs.CV cs.AI 81%

Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments

Laura Alejandra Encinar Gonzalez, John Folkesson, Rudolph Triebel, Riccardo Giubilato

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Under review for ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22900 2025-11-10 cs.CV cs.CL 81%

MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering

Mai A. Shaaban, Tausifa Jan Saleem, Vijay Ram Papineni, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Department of Mathematics and Computer Science, Faculty of Science, Alexandria University(亚历山大大学数学与计算机科学系) Sheikh Shakhbout Medical City(谢赫·沙赫布OUT医疗城)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15252 2025-11-10 cs.CR cs.CL cs.IR 70%

Retrieval-Augmented Review Generation for Poisoning Recommender Systems

Shiyi Yang, Xinshu Li, Guanglin Zhou, Chen Wang, Xiwei Xu, Liming Zhu, Lina Yao

机构 * University of New South Wales and CSIRO’s Data61(新南威尔士大学和CSIRO的Data61)

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05363 2025-11-10 cs.CY cs.AI 57%

AI Literacy for Community Colleges: Instructors' Perspectives on Scenario-Based and Interactive Approaches to Teaching AI

Aparna Maya Warrier, Arav Agarwal, Jaromir Savelka, Christopher A Bogart, Heather Burte

机构 * Computer Science(计算机科学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2504.14245 2025-11-10 cs.CV cs.CL 84%

Towards Explainable Fake Image Detection with Multi-Modal Large Language Models

Yikun Ji, Yan Hong, Jiahui Zhan, Haoxing Chen, jun lan, Huijia Zhu, Weiqiang Wang, Liqing Zhang, Jianfu Zhang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Accepted to ACM MM 2025; 14 pages including Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04831 2025-11-10 cs.RO cs.AI 79%

Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning

NVIDIA, :, Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano-Muñoz, Xinjie Yao, René Zurbrügg, Nikita Rudin, Lukasz Wawrzyniak, Milad Rakhsha, Alain Denzler, Eric Heiden, Ales Borovicka, Ossama Ahmed, Iretiayo Akinola, Abrar Anwar, Mark T. Carlson, Ji Yuan Feng, Animesh Garg, Renato Gasoto, Lionel Gulich, Yijie Guo, M. Gussert, Alex Hansen, Mihir Kulkarni, Chenran Li, Wei Liu, Viktor Makoviychuk, Grzegorz Malczyk, Hammad Mazhar, Masoud Moghani, Adithyavairavan Murali, Michael Noseworthy, Alexander Poddubny, Nathan Ratliff, Welf Rehberg, Clemens Schwarke, Ritvik Singh, James Latham Smith, Bingjie Tang, Ruchik Thaker, Matthew Trepte, Karl Van Wyk, Fangzhou Yu, Alex Millane, Vikram Ramasamy, Remo Steiner, Sangeeta Subramanian, Clemens Volk, CY Chen, Neel Jawale, Ashwin Varghese Kuruttukulam, Michael A. Lin, Ajay Mandlekar, Karsten Patzwaldt, John Welsh, Huihua Zhao, Fatima Anes, Jean-Francois Lafleche, Nicolas Moënne-Loccoz, Soowan Park, Rob Stepinski, Dirk Van Gelder, Chris Amevor, Jan Carius, Jumyung Chang, Anka He Chen, Pablo de Heras Ciechomski, Gilles Daviet, Mohammad Mohajerani, Julia von Muralt, Viktor Reutskyy, Michael Sauter, Simon Schirm, Eric L. Shi, Pierre Terdiman, Kenny Vilella, Tobias Widmer, Gordon Yeoman, Tiffany Chen, Sergey Grizan, Cathy Li, Lotus Li, Connor Smith, Rafael Wiltz, Kostas Alexis, Yan Chang, David Chu, Linxi "Jim" Fan, Farbod Farshidian, Ankur Handa, Spencer Huang, Marco Hutter, Yashraj Narang, Soha Pouya, Shiwei Sheng, Yuke Zhu, Miles Macklin, Adam Moravanszky, Philipp Reist, Yunrong Guo, David Hoeller, Gavriel State

机构 * NVIDIA

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Code and documentation are available here: https://github.com/isaac-sim/IsaacLab

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06980 2025-11-10 cs.CL 79%

Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings

Jonghyun Lee, Dojun Park, Jiwoo Lee, Hoekeon Choi, Sung-Eun Lee

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Published in IEEE Access

Journal ref IEEE Access, vol. 13, pp. 176751-176769, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04690 2025-11-10 cs.MM cs.CL 73%

Automatización de Informes Geotécnicos para Macizos Rocosos con IA

Christofer Valencia, Alexis Llumigusín, Silvia Alvarez, Abrahan Arias, Christian Mejia-Escobar

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.MM

Comments 17 pages, in Spanish language

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 7 篇

2502.15027 2025-11-10 cs.CL cs.AI cs.CV cs.HC 82%

InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback

Henry Hengyuan Zhao, Wenqi Pei, Yifei Tao, Haiyang Mei, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04948 2025-11-10 cs.CV cs.AI 81%

A benchmark multimodal oro-dental dataset for large vision-language models

Haoxin Lv, Ijazul Haq, Jin Du, Jiaxin Ma, Binnian Zhu, Xiaobing Dang, Chaoan Liang, Ruxu Du, Yingjie Zhang, Muhammad Saqib

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04270 2025-11-10 cs.CV cs.AI 81%

ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts

Sangbum Choi, Kyeongryeol Go, Taewoong Jang

机构 * Superb AI Seoul, South Korea(超霸AI首尔韩国)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04977 2025-11-10 cs.CV cs.MM 62%

GSE: Evaluating Sticker Visual Semantic Similarity via a General Sticker Encoder

Heng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo, Weizhi Ma, Min Zhang

机构 * DCST, Tsinghua University(清华大学信息国家实验室) Quan Cheng Laboratory(钱程实验室) AIR, Tsinghua University(清华大学人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04727 2025-11-10 cs.CV cs.AI cs.LG 62%

IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs

Ali Faraz, Akash, Shaharukh Khan, Raja Kolla, Akshat Patidar, Suranjan Goswami, Abhinav Ravi, Chandra Khatri, Shubham Agarwal

机构 * Krutrim AI OLA Electric

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21069 2025-11-10 cs.CV cs.AI cs.HC 62%

GAITEX: Human motion dataset of impaired gait and rehabilitation exercises using inertial and optical sensors

Andreas Spilz, Heiko Oppel, Jochen Werner, Kathrin Stucke-Straub, Felix Capanni, Michael Munz

机构 * AI for Sensor Data Analytics Research Group(人工智能传感器数据解析研究组) Ulm University of Applied Sciences(乌尔姆应用科学大学) Biomechatronic Research Group(生物机械研究组) Institute of Computer Science(计算机科学研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05106 2025-11-10 cs.CV cs.LG 57%

Early Alzheimer's Disease Detection from Retinal OCT Images: A UK Biobank Study

Yasemin Turkan, F. Boray Tek, M. Serdar Nazlı, Öykü Eren

机构 * Department of Computer Engineering Isik University Istanbul, Turkey(计算机工程系伊斯肯大学) Department of Artificial Intelligence and Data Engineering(人工智能与数据工程系) Data Engineering Istanbul Technical University Istanbul, Turkey(数据工程系伊斯坦布尔技术大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏