arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-17 至 2026-02-17 共收录 21 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 21 篇

2509.18053 2026-02-17 cs.RO 82%

V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts

V2V-GoT: 基于多模态大语言模型和思维图的车对车协同自动驾驶

Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Yu-Chiang Frank Wang, Min-Hung Chen, Stephen F. Smith

机构 * NVIDIA Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract)

AI总结 本文提出V2V-GoT框架,结合多模态大语言模型和图-思维方法,提升车对车协同自动驾驶的感知、预测和规划能力。

Comments Accepted by ICRA 2026 (IEEE International Conference on Robotics and Automation). Project: https://eddyhkchiu.github.io/v2vgot.github.io/ Code: https://github.com/eddyhkchiu/V2V-GoT Dataset: https://huggingface.co/datasets/eddyhkchiu/V2V-GoT-QA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13650 2026-02-17 cs.CV cs.AI cs.CL 82%

KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination

KorMedMCQA-V: 一种用于评估视觉语言模型在韩国医学资格考试中多模态多项选择问答能力的基准测试

Byungjin Choi, Seongsu Bae, Sunjun Kweon, Edward Choi

机构 * Ajou University School of Medicine(阿乔大学医学院) KAIST(韩国科学技术院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 KorMedMCQA-V是一个用于评估视觉语言模型在韩国医学资格考试中多模态多项选择问答能力的基准测试,展示了不同模型在医疗领域中的表现差异。

Comments 17 pages, 2 figures, 6 tables. (Includes appendix.)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14285 2026-02-17 cs.DL cs.AI cs.CL cs.LG 81%

FMMD: A multimodal open peer review dataset based on F1000Research

FMMD:基于F1000Research的多模态开放同行评审数据集

Zhenzhen Zhuang, Yuqing Fu, Jing Zhu, Zhangping Zhou, Jialiang Lin

机构 * School of Computer Science and Engineering, Guangzhou Institute of Science and Technology(计算机科学与工程学院,广州科学与技术研究所) College of Foreign Languages and Cultures, Xiamen University(外语学院,厦门大学) Science and Education Evaluation Lab, Guangzhou Institute of Science and Technology(科学与教育评估实验室,广州科学与技术研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 FMMD是一个基于F1000Research的多模态开放同行评审数据集,旨在通过整合多模态数据和版本特定的评审信息,填补现有数据集在同行评审与稿件演变关系上的空白。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14462 2026-02-17 cs.CV cs.CL 81%

RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding

RAVENEA:多模态检索增强视觉文化理解的基准

Jiaang Li, Yifei Yuan, Wenyan Li, Mohammad Aliannejadi, Daniel Hershcovich, Anders Søgaard, Ivan Vulić, Wenxuan Zhang, Paul Pu Liang, Yang Deng, Serge Belongie

机构 * University of Copenhagen(哥本哈根大学) ETH Zürich(苏黎世联邦理工学院) University of Amsterdam(阿姆斯特丹大学) University of Cambridge(剑桥大学) Massachusetts Institute of Technology(麻省理工学院) Singapore University of Technology and Design(新加坡科技设计大学) Singapore Management University(新加坡管理大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 RAVENEA通过多模态检索增强方法提升视觉文化理解,验证了文化注释对多模态检索和下游任务的增强效果,并揭示了不同国家间性能差异。

Comments ICLR 2026; Project page: https://jiaangli.github.io/ravenea/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06288 2026-02-17 cs.CV 79%

Unsupervised MR-US Multimodal Image Registration with Multilevel Correlation Pyramidal Optimization

无监督的MR-US多模态图像配准与多级相关金字塔优化

Jiazheng Wang, Zeyu Liu, Min Liu, Xiang Chen, Xinyao Yu, Yaonan Wang, Hang Zhang

机构 * School of Artificial Intelligence and Robotics, Hunan University, Changsha, Hunan, China(人工智能与机器人学院,湖南大学,长沙,湖南,中国) National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, Hunan, China(机器人视觉感知与控制技术国家工程研究中心,湖南大学,长沙,湖南,中国) National University of Singapore(新加坡国立大学) Cornell University(康奈尔大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出基于多级相关金字塔优化的无监督多模态图像配准方法,以解决术前与术中多模态图像的配准问题,实现了在ReMIND2Reg任务中的优异表现。

Comments first-place method of ReMIND2Reg Learn2Reg MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16505 2026-02-17 cs.CV 79%

PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies

PRISMM-Bench: 一种基于同行评审的多模态不一致基准

Lukas Selch, Yufang Hou, M. Jehanzeb Mirza, Sivan Doveh, James Glass, Rogerio Feris, Wei Lin

机构 * Johannes Kepler University Linz(约翰内斯·开普勒大学林茨分校) Interdisciplinary Transformation University Austria(跨学科转型大学奥地利) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Stanford University(斯坦福大学) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 PRISMM-Bench是首个基于同行评审标记的多模态不一致基准,通过整理384个不一致点,设计三个任务评估模型跨模态检测和推理能力,揭示了多模态科学推理的挑战。

Comments Accepted at ICLR 2026. Project page https://da-luggas.github.io/prismm-bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23522 2026-02-17 cs.CV cs.LG 79%

OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data

OmniEarth-Bench: 向全面评估地球六大球体及跨球体交互的多模态观测地球数据迈进

Fengxiang Wang, Mingshuo Chen, Xuming He, Yi-Fan Zhang, Yueying Li, Feng Liu, Zijie Guo, Zhenghao Hu, Jiong Wang, Jingyi Xu, Zhangrui Li, Junchao Gong, Di Wang, Fenghua Ling, Ben Fei, Weijia Li, Long Lan, Wenjing Yang

机构 * National University of Defense Technology, China(国防科技大学) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室) Beijing University of Posts and Telecommunications, China(北京邮电大学) Zhejiang University, China(浙江大学) Shanghai Jiao Tong University, China(上海交通大学) Fudan University, China(复旦大学) Sun Yat-sen University, China(中山大学) Nanjing University, China(南京大学) University of Science and Technology of China(中国科学技术大学) Wuhan University, China(武汉大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 OmniEarth-Bench是首个全面评估地球六大球体及跨球体交互的多模态基准测试,通过29,855个标准化注释揭示了地球系统认知能力的系统性差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09980 2026-02-17 cs.CV cs.RO 79%

V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models

V2V-LLM:基于多模态大语言模型的车与车协同自动驾驶

Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Stephen F. Smith, Yu-Chiang Frank Wang, Min-Hung Chen

机构 * NVIDIA Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于多模态大语言模型的V2V-LLM,通过车与车协同感知提升自动驾驶安全性和性能。

Comments Accepted by ICRA 2026 (IEEE International Conference on Robotics and Automation). Project: https://eddyhkchiu.github.io/v2vllm.github.io/ Code: https://github.com/eddyhkchiu/V2V-LLM Dataset: https://huggingface.co/datasets/eddyhkchiu/V2V-GoT-QA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14017 2026-02-17 cs.LG 78%

S2SServiceBench: A Multimodal Benchmark for Last-Mile S2S Climate Services

S2SServiceBench:一个多模态基准用于最后一公里S2S气候服务

Chenyue Li, Wen Deng, Zhuotao Sun, Mengxi Jin, Hanzhe Cui, Han Li, Shentong Li, Man Kit Yu, Ming Long Lai, Yuhao Yang, Mengqian Lu, Binhang Yuan

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Nanjing University of Information Science and Technology(南京信息工程大学) Beijing Normal University(北京师范大学)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 S2SServiceBench是一个多模态基准,用于评估S2S气候服务中最后一公里的可靠性,通过10种服务产品和1000多个评估项目,揭示了多模态大语言模型在不确定性下的决策推理挑战。

Comments 18 pages, 3 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13269 2026-02-17 cs.NI 78%

Modality-Tailored Age of Information for Multimodal Data in Edge Computing Systems

多模态数据在边缘计算系统中的模态定制信息年龄

Ying Liu, Yifan Zhang, Xinyu Wang, Chao Yang, Kandaraj Piamrat, Stephan Sigg, Zheng Changr, Yusheng Ji

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文提出模态定制信息年龄(MAoI)度量标准,用于多模态数据在边缘计算中的资源管理和策略优化,并设计了联合采样卸载优化算法以最小化平均 MAoI。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15011 2026-02-17 cs.HC 71%

TouchFusion: Multimodal Wristband Sensing for Ubiquitous Touch Interactions

TouchFusion: 多模态腕带传感用于无处不在的触控交互

Eric Whitmire, Evan Strasnick, Roger Boldu, Raj Sodhi, Nathan Godwin, Shiu Ng, Andre Levi, Amy Karlson, Ran Tan, Josef Faller, Emrah Adamey, Hanchuan Li, Wolf Kienzle, Hrvoje Benko

专题命中 多模态评测 :multimodal(title)

AI总结 TouchFusion通过多模态腕带传感实现无处不在的触控交互,结合多种传感器技术,支持环境和身体表面的触控检测与上下文自适应界面控制。

Comments 23 pages, 22 figures, accompanying video available at https://youtu.be/0fdCwHu7uaA

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20430 2026-02-17 cs.CL cs.AI cs.CV cs.MA 67%

An Agentic System for Rare Disease Diagnosis with Traceable Reasoning

一种具有可追溯推理的罕见病诊断代理系统

Weike Zhao, Chaoyi Wu, Yanjie Fan, Xiaoman Zhang, Pengcheng Qiu, Yuze Sun, Xiao Zhou, Yanfeng Wang, Xin Sun, Ya Zhang, Yongguo Yu, Kun Sun, Weidi Xie

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 DeepRare是一种基于大型语言模型的多代理系统,通过整合40多种专用工具和最新知识源,为罕见病诊断提供决策支持,实现了透明可追溯的推理链。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05430 2026-02-17 cs.CL cs.AI cs.IR cs.LG 62%

ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering

ArtistMus: 一个全球多样、以艺术家为中心的基准,用于检索增强的音乐问答

Daeyong Kwon, SeungHeon Doh, Juhan Nam

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 ArtistMus提出一个全球多样、以艺术家为中心的基准,通过检索增强生成技术提升音乐问答的准确性和上下文推理能力。

Comments Accepted to LREC 2026. This work is an evolution of our earlier preprint arXiv:2507.23334

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01112 2026-02-17 cs.CL 57%

EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation

EmoLoom-2B:基于词典弱监督和KV-Off评估的快速基础模型筛选用于情感分类和VAD预测

Zilin Li, Weiwei Xu, Xuanbo Lu, Zheda Liu

机构 * Zheda Liu(2 刘智达)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 EmoLoom-2B通过词典弱监督和KV-Off评估,快速筛选出适用于情感分类和VAD预测的基础模型。

Comments This paper presents an initial and self-contained study of a lightweight screening pipeline for emotion-aware language modeling, intended as a reproducible baseline and system-level design reference. This latest version corrects and updates certain personal information

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13806 2026-02-17 cs.CV cs.RO 57%

Gaussian Sequences with Multi-Scale Dynamics for 4D Reconstruction from Monocular Casual Videos

具有多尺度动态的高斯序列用于从单目随意视频中进行4D重建

Can Li, Jie Gu, Jingmin Chen, Fangzhou Qiu, Lei Sun

机构 * Nankai University(南开大学) Rightly Robotics

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种基于多尺度动态的高斯序列方法,用于从单目随意视频中实现准确且一致的4D重建,通过多级运动组合和多模态先验约束提升重建保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13771 2026-02-17 cs.MM 57%

SRA: Semantic Relation-Aware Flowchart Question Answering

SRA: 语义关系感知的流程图问答

Xinyu Li, Bowei Zou, Yuchong Chen, Yifan Fan, Yu Hong

专题命中 多模态评测 :multi-modal(abstract);分类 cs.MM

AI总结 SRA通过利用大型语言模型检测节点间的语义关系,提升流程图问答的推理深度和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00220 2026-02-17 eess.IV cs.CV 57%

Deep learning Based Correction Algorithms for 3D Medical Reconstruction in Computed Tomography and Macroscopic Imaging

基于深度学习的3D医学重建在计算机断层扫描和宏观成像中的校正算法

Tomasz Les, Tomasz Markiewicz, Malgorzata Lorent, Miroslaw Dziekiewicz, Krzysztof Siwek

机构 * University of Technology(技术大学) Military Institute of Medicine(军事医学研究院) Institute of Tuberculosis and Lung Diseases(肺结核和肺病研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种混合两阶段配准框架,结合几何先验和深度学习,提升3D医学重建的精度和解剖真实感。

Comments 23 pages, 9 figures, submitted to Applied Sciences (MDPI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01055 2026-02-17 cs.LG cs.AI q-bio.BM q-bio.QM 57%

FGBench: A Dataset and Benchmark for Molecular Property Reasoning at Functional Group-Level in Large Language Models

FGBench: 一个用于大语言模型中功能基团级分子属性推理的数据集和基准

Xuan Liu, Siru Ouyang, Xianrui Zhong, Jiawei Han, Huimin Zhao

机构 * Department of Chemical and Biomolecular Engineering, University of Illinois Urbana-Champaign(化学与生物分子工程系,伊利诺伊大学厄巴纳-香槟分校) Department of Computer Science, University of Illinois Urbana-Champaign(计算机科学系,伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 FGBench通过构建包含功能基团级信息的数据集,提升大语言模型在分子属性推理任务中的能力。

Comments NeurIPS 2025 (Datasets and Benchmarks Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19666 2026-02-17 cs.CL 57%

RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams

RoD-TAL:罗马尼亚驾照考试问答的基准测试

Andrei Vlad Man, Răzvan-Alexandru Smădu, Cristian-George Craciun, Dumitru-Clementin Cercel, Florin Pop, Mihaela-Claudia Cercel

机构 * National University of Science and Technology POLITEHNICA Bucharest, Faculty of Automatic Control and Computers(波兰技术大学布加勒斯特分校) Technical University of Munich(慕尼黑技术大学) National Institute for Research & Development in Informatics - ICI Bucharest(信息研究所-布加勒斯特) Paris 1 Panthéon-Sorbonne University(巴黎1大学) University of Bucharest(布加勒斯特大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 RoD-TAL是一个用于评估大型语言模型和视觉语言模型在罗马尼亚驾照法律问答中性能的多模态基准数据集,通过文本和图像问答任务验证了领域微调和推理优化对考试通过率的影响。

Comments 41 pages, 30 figures, Accepted by the Findings of EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14061 2026-02-17 stat.CO cs.NA math.NA 50%

MPL-HMC: A Tunable Parameterized Leapfrog Framework for Robust Hamiltonian Monte Carlo

MPL-HMC:一种可调参数化Leapfrog框架用于鲁棒哈密顿蒙特卡洛

Sourabh Bhattacharya

专题命中 多模态评测 :multimodal(abstract)

AI总结 MPL-HMC通过可调参数改进HMC,实现鲁棒采样和性能提升,适用于多模分布和复杂模型。

Comments Feedback welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02410 2026-02-17 cs.LG 50%

OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data

OpenTSLM:用于多变量医学文本和时间序列数据推理的时间序列语言模型

Patrick Langer, Thomas Kaar, Max Rosenblattl, Maxwell A. Xu, Winnie Chow, Martin Maritsch, Robert Jakob, Ning Wang, Juncheng Liu, Aradhana Verma, Brian Han, Daniel Seung Kim, Henry Chubb, Scott Ceresnak, Aydin Zahedivash, Alexander Tarlochan Singh Sandhu, Fatima Rodriguez, Daniel McDuff, Elgar Fleisch, Oliver Aalami, Filipe Barata, Paul Schmiedmayer

机构 * Stanford Mussallem Center for Biodesign(斯坦福 Mussallem 生物设计中心) Centre for Digital Health Interventions(数字健康干预中心) Agentic Systems Lab(代理系统实验室) National University of Singapore(新加坡国立大学) Microsoft(微软) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Google Research(谷歌研究) Stanford University(斯坦福大学) Amazon(亚马逊) Division of Cardiovascular Medicine(心血管医学部) Division of Cardiology(心内科部) Pediatric Cardiology(儿童心内科) University of Washington(华盛顿大学)

专题命中 多模态评测 :multimodal(abstract)

AI总结 OpenTSLM通过整合时间序列作为原生模态,提升对多变量医学文本和时间序列数据的推理能力,其模型在多个任务中均优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏