arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9213 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9213 篇

2501.13772 2026-01-13 cs.SD cs.AI cs.LG cs.MM eess.AS 67%

Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models

Jailbreak-AudioBench: 对大型音频语言模型中 jailbreak 威胁的深入评估与分析

Hao Cheng, Erjia Xiao, Jing Shao, Yichi Wang, Le Yang, Chao Shen, Philip Torr, Jindong Gu, Renjing Xu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Oxford(牛津大学) Xi’an Jiaotong University(西安交通大学) Hong Kong University of Science and Technology(香港科技大学) Northeastern University(东北大学) Beijing University of Technology(北京理工大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI、cs.MM、eess.AS

AI总结 Jailbreak-AudioBench 通过构建工具箱、数据集和基准,深入评估大型音频语言模型中 jailbreak 威胁,并促进安全防护机制的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22200 2026-01-13 cs.RO 67%

EnvoDat: A Large-Scale Multisensory Dataset for Robotic Spatial Awareness and Semantic Reasoning in Heterogeneous Environments

EnvoDat:一种大规模多感官数据集,用于机器人空间感知和异构环境中的语义推理

Linus Nwankwo, Bjoern Ellensohn, Vedant Dave, Peter Hofer, Jan Forstner, Marlene Villneuve, Robert Galler, Elmar Rueckert

机构 * Chair of Cyber-Physical System, Montanuniversität Leoben, Austria(智能物理系统系,莱布恩矿业大学,奥地利) Theresianische Militarakademie, Austria(特里西亚军事学院,奥地利) Chair of Subsurface Engineering, Montanuniversität Leoben, Austria(地下工程系,莱布恩矿业大学,奥地利)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract)

AI总结 EnvoDat是一个大规模多感官数据集,用于提升机器人在复杂异构环境中的空间感知和语义推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06106 2026-01-13 cs.LG cs.AI cs.CL cs.CV cs.MA 67%

Judge Model for Large-scale Multimodality Benchmarks

大规模多模态基准的判断模型

Min-Han Shih, Yu-Hsin Wu, Yu-Wei Chen

机构 * Department of Electrical and Computer Engineering, Viterbi School of Engineering, University of Southern California(电气与计算机工程系,维特比工程学院,南加州大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出了一种多模态判断模型,用于评估多种任务,通过聚合多模态判断并生成诊断反馈,展示了其在多模态AI研究中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04897 2026-01-09 cs.CL cs.CV cs.LG cs.MM 67%

V-FAT: Benchmarking Visual Fidelity Against Text-bias

V-FAT:基于文本偏见的视觉真实性基准测试

Ziteng Wang, Yujie He, Guanliang Li, Siqi Yang, Jiaqi Xiong, Songxiang Liu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Meituan(美团) University of Oxford(牛津大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 V-FAT通过三级评估框架量化文本偏见对视觉真实性的干扰,揭示多模态大语言模型在高语言主导性下的视觉崩溃问题。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03733 2026-01-08 cs.CV cs.AI cs.CL cs.CY cs.LG 67%

RadDiff: Describing Differences in Radiology Image Sets with Natural Language

RadDiff:用自然语言描述放射学图像集的差异

Xiaoxian Shen, Yuhui Zhang, Sahithi Ankireddy, Xiaohan Wang, Maya Varma, Henry Guo, Curtis Langlotz, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 RadDiff通过多模态代理系统实现放射学图像集差异的自然语言描述,结合医学知识和多模态推理,在放射学研究配对中取得高准确率,推动临床影像分析的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03594 2026-01-08 cs.CR 67%

Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense

突破大语言模型与视觉语言模型:机制、评估与统一防御

Zejian Chen, Chaozhuo Li, Chao Li, Xi Zhang, Litian Zhang, Yiming He

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract)

AI总结 本文提出三维框架系统回顾大语言模型和视觉语言模型的突破攻击与防御机制,提出统一防御原则并讨论未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24340 2026-01-01 cs.CV cs.AI cs.CL 67%

DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images

DermaVQA-DAS:皮肤科评估方案(DAS)及用于患者生成皮肤科图像中封闭式问答与分割的数据库

Wen-wai Yim, Yujuan Fu, Asma Ben Abacha, Meliha Yetisgen, Noel Codella, Roberto Andres Novoa, Josep Malvehy

机构 * Microsoft Health AI(微软健康人工智能) University of Washington(华盛顿大学) Stanford University(斯坦福大学) Hospital Clinic of Barcelona(巴塞罗那医院诊所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 DermaVQA-DAS引入了皮肤科评估方案DAS,支持封闭式问答与分割任务,通过专家标注数据集和多模态模型评估,提升患者为中心的皮肤科视觉语言模型研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23739 2026-01-01 cs.CL cs.AI cs.CV cs.RO 67%

Break Out the Silverware -- Semantic Understanding of Stored Household Items

拿出银器——对存储家居物品的语义理解

Michaela Levi-Richter, Reuth Mirsky, Oren Glickman

机构 * Bar Ilan University, Computer Science Department(巴伊兰大学计算机科学系) Tufts University, Department of Computer Science(塔夫茨大学计算机科学系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本研究提出NOAM模型,通过结合结构化场景理解和大语言模型推理,实现对家庭物品存储位置的语义理解,显著提升预测准确率并接近人类水平。

Comments Poster presented at the Israeli Seminar on Computational Linguistics 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22257 2025-12-30 q-bio.QM 67%

LiveProteinBench: A Contamination-Free Benchmark for Assessing Models' Specialized Capabilities in Protein Science

LiveProteinBench: 一种无污染的基准测试,用于评估模型在蛋白质科学中的专用能力

Dingyi Rong, Zijian Chen, Qi Jia, Kaiwei Zhang, Haotian Lu, Guangtao Zhai, Ning Liu

专题命中 多模态评测 :multimodal(abstract);multimodal foundation model(abstract)

AI总结 LiveProteinBench通过无污染的多模态基准测试,评估LLM在蛋白质科学中的专用能力,揭示了通用模型在零样本性能上的优势及多模态信息融合的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17319 2025-12-22 cs.CV cs.AI cs.MM 67%

A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs

超高分辨率遥感多模态大语言模型基准

Yunkai Dang, Meiyi Zhu, Donghao Wang, Yizhuo Zhang, Jiacheng Yang, Qi Fan, Yuekun Yang, Wenbin Li, Feng Miao, Yang Gao

机构 * School of Artificial Intelligence Science and Technology, Nanjing University(人工智能科学与技术学院,南京大学) School of Physics, Nanjing University(物理学院,南京大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出RSHR-Bench,一个超高分辨率遥感多模态大语言模型基准,通过高分辨率图像和多样化任务评估视觉理解与推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14496 2025-12-17 cs.CE 67%

BridgeNet: A Dataset of Graph-based Bridge Structural Models for Machine Learning Applications

BridgeNet: 一种基于图的桥梁结构模型数据集用于机器学习应用

Lazlo Bleker, Mustafa Cem Güneş, Pierluigi D'Acunto

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract)

AI总结 BridgeNet是一个包含20,000个桥梁结构模型的公开数据集,用于促进图机器学习和多模态学习在概念性结构设计中的应用。

Comments GNI Symposium on Artifical Intelligence for the Built World 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11074 2025-12-15 cs.CL cs.AI cs.LG cs.MM 67%

MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data

MultiScript30k:利用多语言嵌入扩展跨脚本平行数据

Christopher Driggers-Ellis, Detravious Brinkley, Ray Chen, Aashish Dhawan, Daisy Zhe Wang, Christan Grant

机构 * Computer and Information Science and Engineering(计算机与信息科学与工程)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 MultiScript30k通过多语言嵌入扩展Multi30k数据集,支持多种脚本的全球语言,提升多模态机器翻译的多样性与覆盖范围。

Comments 7 pages, 2 figures, 5 tables. Not published at any conference at this time

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02719 2025-12-03 cs.CL cs.AI cs.CV cs.LG q-bio.NC 67%

Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs

涌现的贝叶斯行为与LLMs中的最优线索整合

Julian Ma, Jun Wang, Zafeirios Fountas

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) AI Centre, Department of Computer Science, University College London(人工智能中心,计算机科学系,伦敦大学学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究揭示LLMs在多模态整合中表现出贝叶斯策略,但准确性不等于鲁棒性,提出贝叶斯一致性分数以评估不确定性处理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00259 2025-12-02 cs.NI 67%

Design and Evaluation of a Multi-Agent Perception System for Autonomous Flying Networks

自主飞行网络多智能体感知系统的設計與評估

Diogo Ferreira, Pedro Ribeiro, André Coelho, Rui Campos

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract)

AI总结 本文提出多智能体感知系统MAPS,利用多模态大语言模型和代理AI,实现自主飞行网络中用户感知与服务需求的自动识别与规范生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21735 2025-12-01 cs.CL cs.AI cs.CV 67%

Closing the Performance Gap Between AI and Radiologists in Chest X-Ray Reporting

弥合AI与放射科医生在胸部X光报告中的性能差距

Harshita Sharma, Maxwell C. Reynolds, Valentina Salvatelli, Anne-Marie G. Sykes, Kelly K. Horst, Anton Schwaighofer, Maximilian Ilse, Olesya Melnichenko, Sam Bond-Taylor, Fernando Pérez-García, Vamshi K. Mugu, Alex Chan, Ceylan Colak, Shelby A. Swartz, Motassem B. Nashawaty, Austin J. Gonzalez, Heather A. Ouellette, Selnur B. Erdal, Beth A. Schueler, Maria T. Wetscherek, Noel Codella, Mohit Jain, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Stephanie Hyland, Panos Korfiatis, Ashish Khandelwal, Javier Alvarez-Valle

机构 * Microsoft(微软公司) Mayo Clinic(梅奥诊所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MAIRA-X通过多模态AI模型在胸部X光报告生成中提升词汇质量、临床正确性和L&T准确性,有效辅助放射科医生,尤其在高患者量的临床环境中

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08438 2025-11-25 cs.CL cs.MM cs.SD eess.AS 67%

CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework

CommonVoice-SpeechRE 和 RPG-MoGe:通过新数据集和多阶生成框架推进语音关系抽取

Jinzhong Ning, Paerhati Tulajiang, Yingying Le, Yijia Zhang, Yuanyuan Sun, Hongfei Lin, Haifeng Liu

机构 * School of Information Science and Technology, Dalian Maritime University(信息科学与技术学院,大连海事大学) School of Computer Science and Technology, Dalian University of Technology(计算机科学与技术学院,大连理工大学) College of Computer Science and Technology, Xinjiang Normal University(计算机科学与技术学院,新疆师范大学) School of Computer and Electronic Information, Nanjing Normal University(计算机与电子信息学院,南京师范大学) Adolescent Education and Intelligence Support Lab of Nanjing Normal University, Laboratory of Philosophy and Social Sciences at Universities in Jiangsu Province(南京师范大学青少年教育与智能支持实验室,江苏省高校哲学社会科学实验室)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CL、cs.MM、eess.AS

AI总结 本文提出CommonVoice-SpeechRE数据集和RPG-MoGe框架,通过多阶生成策略提升语音关系抽取的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22995 2025-11-19 cs.CV cs.AI cs.CL cs.LG 67%

VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning

Jingkun Ma, Runzhe Zhan, Yang Li, Di Sun, Hou Pong Chan, Lidia S. Chao, Derek F. Wong

机构 * NLP(自然语言处理) CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系计算机视觉实验室,澳门大学) University of Macau(澳门大学) DAMO Academy, Alibaba Group(阿里集团达摩院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 58 pages, 28 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22340 2025-11-12 cs.AI cs.CL cs.CV cs.LG 67%

DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry

Changti Wu, Shijie Lian, Zihao Liu, Lei Zhang, Laurence Tianruo Yang, Kai Chen

机构 * East China Normal University(华东师范大学) Zhongguancun Academy(中关村学院) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学) Zhengzhou University(郑州大学) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The code and dataset are available at \href{https://zgca-ai4edu.github.io/DynaSolidGeo/}{DynaSolidGeo}

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10514 2025-11-11 cs.CV cs.AI cs.CL cs.LG 67%

ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness

Yijun Liang, Ming Li, Chenrui Fan, Ziyue Li, Dang Nguyen, Kwesi Cobbina, Shweta Bhardwaj, Jiuhai Chen, Fuxiao Liu, Tianyi Zhou

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS2025. 36 pages, including references and appendix. Code is available at https://github.com/tianyi-lab/ColorBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08221 2025-11-04 cs.CV cs.AI cs.MM 67%

EgoBlind: Towards Egocentric Visual Assistance for the Blind

Junbin Xiao, Nanxin Huang, Hao Qiu, Zhulin Tao, Xun Yang, Richang Hong, Meng Wang, Angela Yao

机构 * National University of Singapore(新加坡国立大学) Communication University of China(中国传媒大学) University of Science and Technology of China(中国科学技术大学) Hefei University of Technoloy(合肥工业大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments NeurIPS'25 (D&B Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17050 2025-11-04 cs.CL cs.AI cs.CE cs.CY cs.MM 67%

Towards Robust Evaluation of STEM Education: Leveraging MLLMs in Project-Based Learning

Xinyi Wu, Yanhao Jia, Qinglin Zhang, Yiran Qin, Luwei Xiao, Shuai Zhao

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22571 2025-10-28 cs.CV cs.AI cs.MM 67%

STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models

Mahiro Ukai, Shuhei Kurita, Nakamasa Inoue

机构 * Institute of Science Tokyo(东京科学研究所) National Institute of Informatics(日本信息处理学会)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10281 2025-10-28 cs.CR cs.AI cs.CL cs.CV cs.LG 67%

ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test

Guan-Yan Yang, Tzu-Yu Cheng, Ya-Wen Teng, Farn Wanga, Kuo-Hui Yeh

机构 * Department of Electrical Engineering, National Taiwan University(国立台湾大学电子工程系) GARMIN (ASIA) CORPORATION(GARMIN(亚洲)公司) Institute of Artificial Intelligence Innovation, National Yang Ming Chiao Tung University(国家阳明交通大学人工智能创新研究所) Department of Information Management, National Dong Hwa University(国立东吴大学资讯管理系)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 30 pages, 22 figures. This preprint has been accepted for publication in Elsevier JOURNAL OF NETWORK AND COMPUTER APPLICATIONS (JNCA)

Journal ref Journal of Network and Computer Applications, Vol. 244, (2025) 104356

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04192 2025-10-14 cs.CV cs.AI cs.CL 67%

ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos

Trinh T. L. Vuong, Jin Tae Kwak

机构 * Korea University(韩国大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09302 2025-10-13 cs.CV cs.AI cs.CL 67%

CapGeo: A Caption-Assisted Approach to Geometric Reasoning

Yuying Li, Siyi Qian, Hao Liang, Leqi Zheng, Ruichuan An, Yongzhen Guo, Wentao Zhang

机构 * THU(清华大学) PKU(北京大学) Ant Group(蚂蚁集团)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments preprint, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06743 2025-10-09 cs.CV cs.AI cs.CL 67%

Evaluating LLMs for Historical Document OCR: A Methodological Framework for Digital Humanities

Maria Levchenko

机构 * Italian Institute of Germanic Studies (IISG)(意大利德语研究学院) University of Bologna(博洛尼亚大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The First Workshop on Natural Language Processing and Language Models for Digital Humanities (LM4DH 2025). RANLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25818 2025-10-01 cs.CV cs.AI cs.CL 67%

VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

Kazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki, Komei Sugiura

机构 * Keio University(庆应大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03214 2025-10-01 cs.CL cs.AI cs.CV 67%

iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs

Julius Mayer, Mohamad Ballout, Serwan Jassim, Farbod Nosrat Nezami, Elia Bruni

机构 * Institute of Cognitive Science, Osnabrück University(认知科学研究所,奥斯纳布吕克大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22991 2025-09-30 cs.CL cs.AI cs.CV cs.IR cs.LG 67%

ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning

Jasin Cekinmez, Omid Ghahroodi, Saad Fowad Chandle, Dhiman Gupta, Ehsaneddin Asgari

机构 * Qatar Computing Research Institute(卡塔尔计算研究所) Princeton University(普林斯顿大学) Virginia Tech(弗吉尼亚理工大学) Amity University(阿米蒂大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07969 2025-09-10 cs.CV cs.AI cs.CL 67%

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

Xin Lai, Junyi Li, Wei Li, Tao Liu, Tianjian Li, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Code, datasets, models are available at https://github.com/Mini-o3/Mini-o3. Project Page: https://mini-o3.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏