arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-03 至 2025-11-03 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2510.27508 2025-11-03 cs.CV cs.AI 84%

Context-Gated Cross-Modal Perception with Visual Mamba for PET-CT Lung Tumor Segmentation

Elena Mulero Ayllón, Linlin Shen, Pierangelo Veltri, Fabrizia Gelardi, Arturo Chiti, Paolo Soda, Matteo Tortora

机构 * Unit of Artificial Intelligence and Computer Systems, Università Campus Bio-Medico di Roma, Italy(人工智能与计算机系统单位,罗马大学生物医学校园,意大利) College of Computer Science and Software Engineering, Shenzhen University, China(计算机科学与软件工程学院,深圳大学,中国) Dept. of Computer Engineering, Modeling, Electronic and System Engineering, University of Calabria, Italy(计算机工程、建模、电子与系统工程系,卡拉布里亚大学,意大利) IRCCS San Raffaele Hospital, Italy(圣拉斐拉医院,意大利) Faculty of Medicine, Vita-Salute San Raffaele University, Italy(医学学院,维塔-桑拉斐拉大学,意大利) Dept. of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering, Umeå University, Sweden(诊断与干预系,辐射物理,生物医学工程,乌梅大学,瑞典) Dept. of Naval, Electrical, Electronics and Telecommunications Engineering, University of Genoa, Italy(海军、电子、电子与电信工程系,热那亚大学,意大利)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27196 2025-11-03 cs.CL cs.AI 84%

MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models

Zixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo, Yayue Deng, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing University of Posts and Telecommunications(北京邮电大学) National University of Singapore(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06594 2025-11-03 cs.CL cs.CV 81%

Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation

Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen

机构 * IRIT, University of Toulouse, France(IRIT,图卢兹大学,法国) Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡) CNRS, IRIT, France(CNRS,IRIT,法国)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27481 2025-11-03 cs.CV 79%

NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding

Wei Xu, Cheng Wang, Dingkang Liang, Zongchuang Zhao, Xingyu Jiang, Peng Zhang, Xiang Bai

机构 * National University of Defense Technology(国防科技大学)

专题命中 多模态评测 :multimodal(title);image-text(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025. Data and models are available at https://github.com/H-EmbodVis/NAUTILUS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26824 2025-11-03 cs.DL cs.AI cs.IR 79%

LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature

Magdalena Lederbauer, Siddharth Betala, Xiyao Li, Ayush Jain, Amine Sehaba, Georgia Channing, Grégoire Germain, Anamaria Leonescu, Faris Flaifil, Alfonso Amayuelas, Alexandre Nozadze, Stefan P. Schmid, Mohd Zaki, Sudheesh Kumar Ethirajan, Elton Pan, Mathilde Franckel, Alexandre Duval, N. M. Anoop Krishnan, Samuel P. Gleason

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

Comments 29 pages, 13 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24003 2025-11-03 cs.LG cs.AI 79%

Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting

ChengAo Shen, Wenchao Yu, Ziming Zhao, Dongjin Song, Wei Cheng, Haifeng Chen, Jingchao Ni

机构 * University of Houston(德克萨斯大学休斯顿分校) NEC Laboratories America(日本 NEC 美国实验室) University of Connecticut(康涅狄格大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27033 2025-11-03 cs.RO cs.AI cs.CV 76%

A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics

Simindokht Jahangard, Mehrzad Mohammadi, Abhinav Dhall, Hamid Rezatofighi

机构 * Faculty of Information Technology, Monash University(信息技术学院,墨尔本大学) Department of Electrical Engineering, Sharif University of Technology(电气工程系,谢赫大学)

专题命中 多模态评测 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27094 2025-11-03 cs.AI 74%

CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning

Hamed Mahdavi, Pouria Mahdavinia, Alireza Farhadi, Pegah Mohammadipour, Samira Malek, Majid Daliri, Pedram Mohammadipour, Alireza Hashemi, Amir Khasahmadi, Vasant Honavar

专题命中 多模态评测 :multimodal(title);分类 cs.AI

Comments Code/data: https://github.com/ref-grader/ref-grader, https://huggingface.co/datasets/combviz/inoi

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26616 2025-11-03 cs.LG cs.AI 57%

Aeolus: A Multi-structural Flight Delay Dataset

Lin Xu, Xinyun Yuan, Yuxuan Liang, Suwan Yin, Yuankai Wu

机构 * Sichuan University(四川大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25761 2025-11-03 cs.CL 57%

DiagramEval: Evaluating LLM-Generated Diagrams via Graphs

Chumeng Liang, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16873 2025-11-03 cs.CV 57%

$\mathtt{M^3VIR}$: A Large-Scale Multi-Modality Multi-View Synthesized Benchmark Dataset for Image Restoration and Content Creation

Yuanzhi Li, Lebin Zhou, Nam Ling, Zhenghao Chen, Wei Wang, Wei Jiang

机构 * Santa Clara University(圣克拉拉大学) University of Newcastle(新castle大学) Futurewei Technologies, Inc.(未来科技公司)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏