arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-22 至 2025-08-22 共收录 42 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 3 篇

2508.15297 2025-08-22 cs.CV cs.AI 81%

DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding

Zhu Wang, Homaira Huda Shomee, Sathya N. Ravi, Sourav Medya

机构 * Department of Computer Science, University of Illinois Chicago(计算机科学系,伊利诺伊大学芝加哥分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by EMNLP 2025. 22 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00329 2025-08-22 cs.CV cs.LG 79%

ABC: Achieving Better Control of Multimodal Embeddings using VLMs

Benjamin Schneider, Florian Kerschbaum, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10225 2025-08-22 cs.CV 57%

Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection

Jinglun Li, Kaixun Jiang, Zhaoyu Chen, Bo Lin, Yao Tang, Weifeng Ge, Wenqiang Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai(智能机器人与先进制造学院,复旦大学,上海) Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University, Shanghai(上海智能信息处理重点实验室,计算机科学与人工智能学院,复旦大学,上海) JIIOV Technology, Beijing(JIIOV技术,北京)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2508.12227 2025-08-22 cs.CL 79%

Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges

Abdelhamid Haouhat, Slimane Bellaouar, Attia Nehar, Hadda Cherroun, Ahmed Abdelali

机构 * Ziane Achour University(赞赞·阿赫尔大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15565 2025-08-22 cs.SD 78%

Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization

Liping Chen, Chenyang Guo, Rui Wang, Kong Aik Lee, Zhenhua Ling

机构 * University of Science and Technology of China(中国科学技术大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 音频语音多模态 :any-to-any(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14976 2025-08-22 cs.LG 78%

Aura-CAPTCHA: A Reinforcement Learning and GAN-Enhanced Multi-Modal CAPTCHA System

Joydeep Chandra, Prabal Manhas, Ramanjot Kaur, Rashi Sahay

机构 * Department of Computer Science and Engineering, Chandigarh University, Mohali, Punjab, India(昌迪加尔大学计算机科学与工程系,莫哈利,旁遮普,印度) Department of Computer Science and Engineering, Manav Rachna International Institute of Research and Studies(曼纳瓦拉国际研究与学习研究所计算机科学与工程系)

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14912 2025-08-22 cs.IR 78%

Multimodal Recommendation via Self-Corrective Preference Alignmen

Yalong Guan, Xiang Chen, Mingyang Wang, Xiangyu Wu, Lihao Liu, Chao Qi, Shuang Yang, Tingting Gao, Guorui Zhou, Changjian Chen

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15407 2025-08-22 cs.CL cs.AI 73%

When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models

Cheng Wang, Gelei Deng, Xianglin Yang, Han Qiu, Tianwei Zhang

机构 * National University of Singapore(国立新加坡大学) Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12918 2025-08-22 cs.SD 50%

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

Lei Zhao, Rujin Chen, Chi Zhang, Xiao-Lei Zhang, Xuelong Li

机构 * School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Research and Development Institute of Northwestern Polytechnical University in Shenzhen, China(西北工业大学深圳研发院,中国)

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2508.15717 2025-08-22 cs.CV cs.AI 62%

StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Yanlai Yang, Zhuokai Zhao, Satya Narayan Shukla, Aashu Singh, Shlok Kumar Mishra, Lizhu Zhang, Mengye Ren

机构 * Meta AI New York University(纽约大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14941 2025-08-22 cs.MM cs.CL 62%

Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.MM

Comments 12 pages, 4 figures, 2 tables. Extends our earlier framework on hierarchical narrative graphs with a semantic normalization module

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11988 2025-08-22 cs.CV 57%

Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis

Nicolas Mastropasqua, Ignacio Bugueno-Cordova, Rodrigo Verschae, Daniel Acevedo, Pablo Negri, Maria E. Buemi

机构 * Universidad de Buenos Aires, Facultad de Ciencias Exactas y Naturales(布宜诺斯艾利斯大学,精确科学与自然学院) Institute of Engineering Sciences, Universidad de O’Higgins(工程科学研究所,奥希金斯大学) CONICET-UBA, Instituto de Ciencias de la Computacion (ICC)(CONICET-UBA,计算科学研究所) L3S Research Center, Leibniz Universität Hannover(L3S研究中心,汉诺威莱布尼茨大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); 2nd Workshop on Neuromorphic Vision (NeVi)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15036 2025-08-22 cs.CR cs.AI 57%

MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs

Ruyi Ding, Tianhong Xu, Xinyi Shen, Aidong Adam Ding, Yunsi Fei

机构 * Louisiana State University(路易斯安那州立大学) Northeastern University(东北大学) Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments This paper will appear in CCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14942 2025-08-22 cs.LG 50%

Structure-Aware Temporal Modeling for Chronic Disease Progression Prediction

Jiacheng Hu, Bo Zhang, Ting Xu, Haifeng Yang, Min Gao

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2501.15183 2025-08-22 cs.IR 78%

Generating Negative Samples for Multi-Modal Recommendation

Yanbiao Ji, Dan Luo, Chang Liu, Shaokai Wu, Jing Tong, Qicheng He, Deyi Ji, Hongtao Lu, Yue Ding

专题命中 跨模态检索 :multi-modal(title,abstract)

Comments Accepted by ACM Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 1 篇

2505.22306 2025-08-22 cs.LG cs.AI 57%

Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer

Zehua Chen, Yuyang Miao, Liyuan Wang, Luyun Fan, Danilo P. Mandic, Jun Zhu

机构 * Department of Computer Science and Technology, Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint Center for ML, Tsinghua University, Beijing, China(计算机科学与技术系、人工智能研究院、BNRist中心、THBI实验室、清华大学-博世联合机器学习中心、清华大学、北京,中国) Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing, China(心理学与认知科学系、清华大学、北京,中国) Department of Electrical and Electronic Engineering, Imperial College London, London, United Kingdom(电子与电气工程系、伦敦帝国理工学院、伦敦,英国) Beijing Anzhen Hospital of Capital Medical University, Beijing Institute of Heart Lung and Blood Vessel Diseases, Chinese Institutes for Medical Research, Beijing, China(首都医科大学北京安贞医院、北京心肺血管疾病研究院、中国医学研究院、北京,中国)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 12 篇

2508.07470 2025-08-22 cs.CV 87%

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

Siminfar Samakoush Galougah, Rishie Raj, Sanjoy Chowdhury, Sayan Nag, Ramani Duraiswami

专题命中 多模态评测 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);omni-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15370 2025-08-22 cs.CL cs.AI 84%

Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation

Yichi Zhang, Yao Huang, Yifan Wang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Huanran Chen, Xiao Yang, Xingxing Wei, Hang Su, Yinpeng Dong, Jun Zhu

机构 * Department of Computer Science and Technology, College of AI, Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(计算机科学与技术系、人工智能学院、人工智能研究所、清华-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学) Institute of Artificial Intelligence, Beihang University(人工智能研究院、北航) RealAI

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments For Appendix, please refer to arXiv:2406.07057

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11060 2025-08-22 cs.CV 83%

BannerAgency: Advertising Banner Design with Multimodal LLM Agents

Heng Wang, Yotaro Shimose, Shingo Takamatsu

机构 * Sony Group Corporation(索尼集团)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11169 2025-08-22 cs.CL cs.AI 81%

MuSeD: A Multimodal Spanish Dataset for Sexism Detection in Social Media Videos

Laura De Grazia, Pol Pastells, Mauro Vázquez Chas, Desmond Elliott, Danae Sánchez Villegas, Mireia Farrús, Mariona Taulé

机构 * University of Barcelona, CLiC-Language and Computing Center(巴塞罗那大学,CLiC语言与计算中心) University of Copenhagen, Department of Computer Science(哥本哈根大学,计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments COLM 2025 camera-ready version: expanded Section 4.3 with an additional experiment using an extended definition-based prompt (including a definition of sexist content), and applied minor corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15299 2025-08-22 cs.CV 79%

BasketLiDAR: The First LiDAR-Camera Multimodal Dataset for Professional Basketball MOT

Ryunosuke Hayashi, Kohei Torimi, Rokuto Nagata, Kazuma Ikeda, Ozora Sako, Taichi Nakamura, Masaki Tani, Yoshimitsu Aoki, Kentaro Yoshioka

机构 * Keio University(Keio大学) AISIN CORPORATION(AISIN公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to MMSports

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15481 2025-08-22 cs.IR 78%

On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking

Fang Wang, Yongjie Wang, Zonghao Yang, Minghao Hu, Xiaoying Bai

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15189 2025-08-22 cs.AI cs.CV eess.IV 62%

SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis

Jiahao Xu, Changchang Yin, Odysseas Chatzipanagiotou, Diamantis Tsilimigras, Kevin Clear, Bingsheng Yao, Dakuo Wang, Timothy Pawlik, Ping Zhang

机构 * The Ohio State University(俄亥俄州立大学) The Ohio State University Wexner Medical Center(俄亥俄州立大学韦克斯纳医学中心) Northeastern University(东北大学)

专题命中 多模态评测 :MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06994 2025-08-22 cs.CV cs.AI 62%

Cross-Modality Masked Learning for Survival Prediction in ICI Treated NSCLC Patients

Qilong Xing, Zikai Song, Bingxin Gong, Lian Yang, Junqing Yu, Wei Yang

机构 * School of Computer Science and Technology(计算机科学与技术学院) Department of Radiology, Union Hospital, Tongji Medical College(放射科、同济医学院附属医院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15353 2025-08-22 cs.CV 57%

RCDINO: Enhancing Radar-Camera 3D Object Detection with DINOv2 Semantic Features

Olga Matykina, Dmitry Yudin

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院) AIRI

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication in Optical Memory and Neural Networks, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15232 2025-08-22 cs.CV 57%

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation

Ruipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang, Shifeng Zhang, Xu Zhou, Liang Wang, Si Liu

机构 * Beihang University(北京航空航天大学) Sangfor Technologies Inc.(深信服科技有限公司) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08324 2025-08-22 cs.AI 57%

ADAM: An AI Reasoning and Bioinformatics Model for Alzheimer's Disease Detection and Microbiome-Clinical Data Integration

Ziyuan Huang, Vishaldeep Kaur Sekhon, Roozbeh Sadeghian, Maria L. Vaida, Cynthia Jo, Doyle Ward, Vanni Bucci, John P. Haran

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14934 2025-08-22 q-bio.GN cs.LG 50%

AGP: A Novel Arabidopsis thaliana Genomics-Phenomics Dataset and its HyperGraph Baseline Benchmarking

Manuel Serna-Aguilera, Fiona L. Goggin, Aranyak Goswami, Alexander Bucksch, Suxing Liu, Khoa Luu

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系) Department of Entomology and Plant Pathology(昆虫学与植物病理学系) Department of Animal Science(动物科学系) School of Plant Sciences(植物科学学院)

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 5 篇

2508.15043 2025-08-22 cs.HC 78%

LitForager: Exploring Multimodal Literature Foraging Strategies in Immersive Sensemaking

Haoyang Yang, Elliott H. Faa, Weijian Liu, Shunan Guo, Duen Horng Chau, Yalong Yang

专题命中 多模态Agent :multimodal(title,abstract)

Comments 11 pages, 10 figures, Accepted to IEEE ISMAR 2025 (TVCG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22146 2025-08-22 cs.CV cs.AI cs.CL q-bio.NC 67%

Flexible Tool Selection through Low-dimensional Attribute Alignment of Vision and Language

Guangfu Hao, Haojie Wen, Liangxuan Guo, Yang Chen, Yanchao Bi, Shan Yu

机构 * Laboratory of Brain Atlas and Brain-inspired Intelligence, Institute of Automation Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所脑图谱与类脑智能实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) School of Systems Science, Beijing Normal University(北京师范大学系统科学学院) School of Psychological and Cognitive Sciences & Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京大学心理与认知科学学院) IDG/McGovern Institute for Brain Research, Peking University(北京大学IDG/ McGovern脑科学研究院) Institute for Artificial Intelligence & Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学人工智能研究所) School of Future Technology, University of Chinese Academy of Sciences (UCAS)(中国科学院大学未来技术学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏