arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-22 至 2025-08-22 共收录 12 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 12 篇

2508.07470 2025-08-22 cs.CV 87%

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

Siminfar Samakoush Galougah, Rishie Raj, Sanjoy Chowdhury, Sayan Nag, Ramani Duraiswami

专题命中 多模态评测 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);omni-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15370 2025-08-22 cs.CL cs.AI 84%

Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation

Yichi Zhang, Yao Huang, Yifan Wang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Huanran Chen, Xiao Yang, Xingxing Wei, Hang Su, Yinpeng Dong, Jun Zhu

机构 * Department of Computer Science and Technology, College of AI, Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(计算机科学与技术系、人工智能学院、人工智能研究所、清华-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学) Institute of Artificial Intelligence, Beihang University(人工智能研究院、北航) RealAI

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments For Appendix, please refer to arXiv:2406.07057

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11060 2025-08-22 cs.CV 83%

BannerAgency: Advertising Banner Design with Multimodal LLM Agents

Heng Wang, Yotaro Shimose, Shingo Takamatsu

机构 * Sony Group Corporation(索尼集团)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11169 2025-08-22 cs.CL cs.AI 81%

MuSeD: A Multimodal Spanish Dataset for Sexism Detection in Social Media Videos

Laura De Grazia, Pol Pastells, Mauro Vázquez Chas, Desmond Elliott, Danae Sánchez Villegas, Mireia Farrús, Mariona Taulé

机构 * University of Barcelona, CLiC-Language and Computing Center(巴塞罗那大学,CLiC语言与计算中心) University of Copenhagen, Department of Computer Science(哥本哈根大学,计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments COLM 2025 camera-ready version: expanded Section 4.3 with an additional experiment using an extended definition-based prompt (including a definition of sexist content), and applied minor corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15299 2025-08-22 cs.CV 79%

BasketLiDAR: The First LiDAR-Camera Multimodal Dataset for Professional Basketball MOT

Ryunosuke Hayashi, Kohei Torimi, Rokuto Nagata, Kazuma Ikeda, Ozora Sako, Taichi Nakamura, Masaki Tani, Yoshimitsu Aoki, Kentaro Yoshioka

机构 * Keio University(Keio大学) AISIN CORPORATION(AISIN公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to MMSports

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15481 2025-08-22 cs.IR 78%

On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking

Fang Wang, Yongjie Wang, Zonghao Yang, Minghao Hu, Xiaoying Bai

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15189 2025-08-22 cs.AI cs.CV eess.IV 62%

SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis

Jiahao Xu, Changchang Yin, Odysseas Chatzipanagiotou, Diamantis Tsilimigras, Kevin Clear, Bingsheng Yao, Dakuo Wang, Timothy Pawlik, Ping Zhang

机构 * The Ohio State University(俄亥俄州立大学) The Ohio State University Wexner Medical Center(俄亥俄州立大学韦克斯纳医学中心) Northeastern University(东北大学)

专题命中 多模态评测 :MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06994 2025-08-22 cs.CV cs.AI 62%

Cross-Modality Masked Learning for Survival Prediction in ICI Treated NSCLC Patients

Qilong Xing, Zikai Song, Bingxin Gong, Lian Yang, Junqing Yu, Wei Yang

机构 * School of Computer Science and Technology(计算机科学与技术学院) Department of Radiology, Union Hospital, Tongji Medical College(放射科、同济医学院附属医院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15353 2025-08-22 cs.CV 57%

RCDINO: Enhancing Radar-Camera 3D Object Detection with DINOv2 Semantic Features

Olga Matykina, Dmitry Yudin

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院) AIRI

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication in Optical Memory and Neural Networks, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15232 2025-08-22 cs.CV 57%

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation

Ruipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang, Shifeng Zhang, Xu Zhou, Liang Wang, Si Liu

机构 * Beihang University(北京航空航天大学) Sangfor Technologies Inc.(深信服科技有限公司) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08324 2025-08-22 cs.AI 57%

ADAM: An AI Reasoning and Bioinformatics Model for Alzheimer's Disease Detection and Microbiome-Clinical Data Integration

Ziyuan Huang, Vishaldeep Kaur Sekhon, Roozbeh Sadeghian, Maria L. Vaida, Cynthia Jo, Doyle Ward, Vanni Bucci, John P. Haran

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14934 2025-08-22 q-bio.GN cs.LG 50%

AGP: A Novel Arabidopsis thaliana Genomics-Phenomics Dataset and its HyperGraph Baseline Benchmarking

Manuel Serna-Aguilera, Fiona L. Goggin, Aranyak Goswami, Alexander Bucksch, Suxing Liu, Khoa Luu

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系) Department of Entomology and Plant Pathology(昆虫学与植物病理学系) Department of Animal Science(动物科学系) School of Plant Sciences(植物科学学院)

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏