arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-27 至 2025-10-27 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2306.13394 2025-10-27 cs.CV 84%

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, Rongrong Ji, Caifeng Shan, Ran He

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) Tencent Youtu Lab(腾讯优图实验室) Xiamen University(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS DB 2025 Spotlight, Project Page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21603 2025-10-27 cs.IR cs.CL 83%

Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research

Kuicai Dong, Shurui Huang, Fangda Ye, Wei Han, Zhi Zhang, Dexun Li, Wenjun Li, Qu Yang, Gang Wang, Yichao Wang, Chen Zhang, Yong Liu

机构 * Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21111 2025-10-27 cs.CV 83%

PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments

Weijie Zhou, Xuantang Xiong, Yi Peng, Manli Tao, Chaoyang Zhao, Honghui Dong, Ming Tang, Jinqiao Wang

机构 * Beijing Jiaotong University(北京交通大学) Tencent Robotics X & Futian Laboratory, Shenzhen(腾讯机器人X及福田实验室,深圳) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心,中国科学院自动化研究所) ObjectEye Inc(ObjectEye公司)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments 39th Conference on Neural Information Processing Systemss (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21438 2025-10-27 cs.RO 82%

PREVENT: Proactive Risk Evaluation and Vigilant Execution of Tasks for Mobile Robotic Chemists using Multi-Modal Behavior Trees

Satheeshkumar Veeramani, Zhengxue Zhou, Francisco Munguia-Galeano, Hatem Fakhruldeen, Thomas Roddelkopf, Mohammed Faeik Ruzaij Al-Okby, Kerstin Thurow, Andrew Ian Cooper

机构 * Department of Chemistry and Material Innovation Factory, University of Liverpool(化学系和材料创新工厂,利物浦大学) Center for Life Science Automation, University of Rostock(生命科学自动化中心,罗斯托克大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments 25 pages, 8 figures, paper submitted to Robotics and Autonomous Systems Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21679 2025-10-27 cs.AI 79%

A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection

Gaku Morio, Harri Rowlands, Dominik Stammbach, Christopher D. Manning, Peter Henderson

机构 * Hitachi, Ltd.(日立公司) Stanford University(斯坦福大学) Centre for the Acceleration of Social Technology(社会技术加速中心) Princeton University(普林斯顿大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Forthcoming in NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11568 2025-10-27 q-bio.QM cs.AI cs.LG 79%

BioCube: A Multimodal Dataset for Biodiversity Research

Stylianos Stasinos, Martino Mensio, Elena Lazovik, Athanasios Trantas

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments published to BiDS'25, 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21214 2025-10-27 cs.CR 78%

Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses

Xingwei Zhong, Kar Wai Fok, Vrizlynn L. L. Thing

专题命中 多模态评测 :MLLM(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20967 2025-10-27 cs.CV cs.AI 62%

3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models

Sraavya Sambara, Sung Eun Kim, Xiaoman Zhang, Luyang Luo, Shreya Johri, Mohammed Baharoon, Du Hyun Ro, Pranav Rajpurkar

机构 * Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院) Seoul National University Hospital(首尔国立大学医院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21160 2025-10-27 cs.CV 57%

Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study

Guanlin Wu, Boyan Su, Yang Zhao, Pu Wang, Yichen Lin, Hao Frank Yang

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12448 2025-10-27 cs.CV 57%

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

Yang Liu, Ming Ma, Xiaomin Yu, Pengxiang Ding, Han Zhao, Mingyang Sun, Siteng Huang, Donglin Wang

机构 * Westlake University(西湖大学) Zhejiang University(浙江大学) Harbin Institute of Technology(哈尔滨工业大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08646 2025-10-27 cs.AI 57%

P-CAFE: Personalized Cost-Aware Incremental Feature Selection For Electronic Health Records

Naama Kashani, Mira Cohen, Uri Shaham

机构 * Bar-Ilan University(巴伊兰大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏