arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 20 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 20 篇

2511.06722 2025-11-11 cs.CV cs.AI cs.CL 85%

Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View

Jianyu Qi, Ding Zou, Wenrui Yan, Rui Ma, Jiaxu Li, Zhijie Zheng, Zhiguo Yang, Rongchang Zhao

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accpeted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19875 2025-11-11 cs.CV 83%

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

Kirolos Ataallah, Eslam Abdelrahman, Mahmoud Ahmed, Chenhui Gou, Khushbu Pahwa, Jian Ding, Mohamed Elhoseiny

机构 * KAUST(卡塔尔科技大学) Monash University(墨尔本大学) RICE University(里士满大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted for oral presentation at the EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07010 2025-11-11 cs.CL cs.CV cs.HC 81%

A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation

Siddharth Betala, Kushan Raj, Vipul Betala, Rohan Saswade

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at The 12th Workshop on Asian Translation, co-located with IJCLNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20665 2025-11-11 cs.CV cs.MM 81%

SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection

Yuxuan Li, Xiang Li, Yunheng Li, Yicheng Zhang, Yimian Dai, Qibin Hou, Ming-Ming Cheng, Jian Yang

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted as Oral in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05883 2025-11-11 cs.AI 79%

Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks

Hehai Lin, Hui Liu, Shilei Cao, Jing Li, Haoliang Li, Wenya Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) City University of Hong Kong(香港城市大学) Sun Yat-sen University(中山大学) Harbin Institute of Technology(哈尔滨工业大学) Nanyang Technological University(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18004 2025-11-11 cs.CV 79%

SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning

Yuhao Shen, Liyuan Sun, Yan Xu, Wenbin Liu, Shuping Zhang, Shawn Afvari, Zhongyi Han, Jiaoyan Song, Yongzhi Ji, Tao Lu, Xiaonan He, Xin Gao, Juexiao Zhou

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK–Shenzhen)(数据科学学院,香港中文大学(深圳)) Computer Science Program, CEMSE Division, King Abdullah University of Science and Technology (KAUST)(计算机科学项目,科学与工程学院,国王 Abdullah 科学技术大学) Center of Excellence on Smart Health, KAUST(智能健康卓越中心,国王 Abdullah 科学技术大学) Center of Excellence for Generative AI, KAUST(生成式人工智能卓越中心,国王 Abdullah 科学技术大学) Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University(皮肤科,北京安贞医院,首都医科大学) Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital(皮肤科,天津整合皮肤科研究院,天津中医药大学附属医院) Department of Dermatology, Beijing Aerospace General Hospital(皮肤科,北京航天总医院) Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院) DermAssure, LLC(DermAssure 公司) School of Medicine, New York Medical College(医学院,纽约医学院) Capital Medical University(首都医科大学) Department of Dermatology, Second Hospital of Jilin University(皮肤科,吉林大学第二医院) Emergency Critical Care Center, Beijing AnZhen Hospital, Capital Medical University(急诊重症中心,北京安贞医院,首都医科大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07276 2025-11-11 cs.LG 78%

RobustA: Robust Anomaly Detection in Multimodal Data

Salem AlMarri, Muhammad Irzam Liaqat, Muhammad Zaigham Zaheer, Shah Nawaz, Karthik Nandakumar, Markus Schedl

机构 * Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) IMT School for Advanced Studies(IMT高级研究学院) Johannes Kepler University Linz(林茨约瑟夫·冯·克莱门茨大学) Human-centered AI Group, AI Lab, Linz Institute of Technology(以人为中心的人工智能小组、人工智能实验室、林茨技术研究所)

专题命中 多模态评测 :multimodal(title,abstract)

Comments Submitted to IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06020 2025-11-11 cs.DB 78%

RF-Behavior: A Multimodal Radio-Frequency Dataset for Human Behavior and Emotion Analysis

Si Zuo, Yuqing Song, Sahar Golipoor, Ying Liu, Xujun Ma, Stephan Sigg

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05596 2025-11-11 cs.LG physics.comp-ph physics.flu-dyn 78%

AutoHood3D: A Multi-Modal Benchmark for Automotive Hood Design and Fluid-Structure Interaction

Vansh Sharma, Harish Jai Ganesh, Maryam Akram, Wanjiao Liu, Venkat Raman

机构 * University of Michigan Ann Arbor(密歇根大学安娜堡分校) Ford Research and Innovation Center(福特研究与创新中心)

专题命中 多模态评测 :multi-modal(title,abstract)

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03248 2025-11-11 cs.CR 76%

Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework

Junhao Li, Jiahao Chen, Zhou Feng, Chunyi Zhou

专题命中 多模态评测 :multimodal(abstract,comments);multi-modal(abstract);cross-modal(abstract)

Comments 14 pages, 3 figures; Accepted by MMM 2026; Complete version in progress. Dataset available at https://huggingface.co/datasets/xaddh/multimodal-privacy

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10534 2025-11-11 cs.CV cs.LG cs.MM 73%

MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates

Binyu Zhao, Wei Zhang, Zhaonian Zou

机构 * The School of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术学院,哈尔滨工业大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments This is the accepted version of an article that has been published in \textbf{Pattern Recognition}. The final version is available via the DOI, or for 50 days' free access via this Share Link: https://authors.elsevier.com/a/1m40D77nKsBm- (valid until December 28, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04910 2025-11-11 cs.CL 70%

SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents

Jaehoon Lee, Sohyun Kim, Wanggeun Park, Geon Lee, Seungkyung Kim, Minyoung Lee

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

Comments 27 pages, 15 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10514 2025-11-11 cs.CV cs.AI cs.CL cs.LG 67%

ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness

Yijun Liang, Ming Li, Chenrui Fan, Ziyue Li, Dang Nguyen, Kwesi Cobbina, Shweta Bhardwaj, Jiuhai Chen, Fuxiao Liu, Tianyi Zhou

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS2025. 36 pages, including references and appendix. Code is available at https://github.com/tianyi-lab/ColorBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05542 2025-11-11 q-bio.NC cs.AI cs.CV cs.LG 62%

ConnectomeBench: Can LLMs Proofread the Connectome?

Jeff Brown, Andrew Kirjner, Annika Vivekananthan, Ed Boyden

机构 * MIT(麻省理工学院) HHMI(霍华德·休斯医学研究所) McGovern Institute(麦戈文研究所) MIT Departments of Brain and Cognitive Sciences(麻省理工学院脑科学与认知科学系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments To appear in NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06840 2025-11-11 cs.CV cs.RO 57%

PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory

Qunchao Jin, Yilin Wu, Changhao Chen

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted as a poster in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06752 2025-11-11 cs.CV 57%

Med-SORA: Symptom to Organ Reasoning in Abdomen CT Images

You-Kyoung Na, Yeong-Jun Cho

机构 * Chonnam National University(全南国立大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06522 2025-11-11 cs.AI cs.LG 57%

FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis

Jan Ondras, Marek Šuppa

机构 * MIT(麻省理工学院) Comenius University in Bratislava(布拉迪斯拉发孔院大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments Accepted to The 5th Workshop on Mathematical Reasoning and AI at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025); 25 pages, 14 figures, 8 tables; Code available at https://github.com/NaiveNeuron/FractalBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04307 2025-11-11 cs.AI 57%

GUI-360$^\circ$: A Comprehensive Dataset and Benchmark for Computer-Using Agents

Jian Mu, Chaoyun Zhang, Chiming Ni, Lu Wang, Bo Qiao, Kartik Mathur, Qianhui Wu, Yuhang Xie, Xiaojun Ma, Mengyu Zhou, Si Qin, Liqun Li, Yu Kang, Minghua Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang

机构 * Nanjing University(南京大学) Microsoft(微软) ZJU-UIUC(浙大-UIUC) Peking University(北京大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00527 2025-11-11 eess.IV cs.CV 57%

MAROON: A Dataset for the Joint Characterization of Near-Field High-Resolution Radio-Frequency and Optical Depth Imaging Techniques

Vanessa Wirth, Johanna Bräunig, Nikolai Hofmann, Martin Vossiek, Tim Weyrich, Marc Stamminger

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(埃朗根-纽伦堡弗里德里希-亚历山大大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06360 2025-11-11 cs.CV 57%

AesTest: Measuring Aesthetic Intelligence from Perception to Production

Guolong Wang, Heng Huang, Zhiqiang Zhang, Wentian Li, Feilong Ma, Xin Jin

机构 * University of International Business and Economics(国际商务经济大学) University of Science and Technology of China(中国科学技术大学) Huawei Technologies Co., Ltd(华为技术有限公司) Beijing Electronic Science and Technology Institute(北京电子科技研究所) Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏