arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-17 至 2025-12-17 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 14 篇

2512.14594 2025-12-17 cs.CV 83%

LLM-driven Knowledge Enhancement for Multimodal Cancer Survival Prediction

基于大语言模型的知识增强的多模态癌症生存预测

Chenyu Zhao, Yingxue Xu, Fengtao Zhou, Yihui Wang, Hao Chen

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) Division of Life Science, The Hong Kong University of Science and Technology(香港科技大学生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出KEMM模型,通过整合专家报告和预后背景知识,利用知识增强的跨模态注意力模块提升癌症生存预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14620 2025-12-17 cs.CL cs.AI cs.CV 82%

JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction

JMMMU-Pro:通过Vibe基准构建的基于图像的日本多学科多模态理解基准

Atsuyuki Miyai, Shota Onohara, Jeonghun Baek, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 JMMMU-Pro通过构建基于图像的多模态理解基准,评估LMMs在日文处理能力,提出Vibe基准构建方法以提高基准质量。

Comments Project page: https://mmmu-japanese-benchmark.github.io/JMMMU_Pro/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14491 2025-12-17 cs.AI 79%

Sparse Multi-Modal Transformer with Masking for Alzheimer's Disease Classification

用于阿尔茨海默病分类的稀疏多模态Transformer与掩码

Cheng-Han Lu, Pei-Hsuan Tsai

机构 * Cheng-Han Lu(无) Pei-Hsuan Tsai(无)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出SMMT,一种稀疏多模态Transformer架构,通过稀疏注意力和模态掩码提升效率与鲁棒性,应用于阿尔茨海默病分类任务。

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01728 2025-12-17 cs.CV 79%

Multimodal classification of forest biodiversity potential from 2D orthophotos and 3D airborne laser scanning point clouds

基于2D正射影像和3D空中激光扫描点云的森林生物多样性潜力多模态分类

Simon B. Jensen, Stefan Oehmcke, Andreas Møgelmose, Meysam Madadi, Christian Igel, Sergio Escalera, Thomas B. Moeslund

机构 * Perception Laboratory, Aalborg University, Denmark Pioneer Centre for Artificial Intelligence, Denmark Department of Computer Science, Copenhagen University, Denmark Institute for Visual \& Analytic Computing, Rostock University, Germany University of Barcelona Computer Vision Center, Spain

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究利用2D正射影像和3D ALS点云数据,通过多模态深度学习融合方法,实现对森林生物多样性潜力的高效评估,实验结果达到82%的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07553 2025-12-17 cs.AI 79%

COMMA: A Communicative Multimodal Multi-Agent Benchmark

COMMA:一种基于通信的多模态多智能体基准

Timothy Ossowski, Danyal Maqbool, Jixuan Chen, Zefan Cai, Tyler Bradshaw, Junjie Hu

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校) Department of Computer Sciences UC San Diego(计算机科学系加州大学圣地亚哥分校) Department of Radiology University of Wisconsin-Madison(放射学系威斯康星大学麦迪逊分校) Department of Computer Sciences Department of Biostatistics and Medical Informatics University of Wisconsin-Madison(计算机科学系生物统计学与医学信息学系威斯康星大学麦迪逊分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 COMMA基准通过语言通信评估多模态多智能体系统的协作性能,揭示现有模型在智能体协作中的不足。

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13573 2025-12-17 cs.CV 79%

MMhops-R1: Multimodal Multi-hop Reasoning

MMhops-R1: 多模态多跳推理

Tao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Chen, Bing Li, Chunfeng Yuan, Guangting Wang, Fengyun Rao, Ying Shan, Weiming Hu

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 MMhops-R1通过动态规划和多模态知识整合,提升多跳推理能力,提出新型mRAG框架并提供挑战性基准。

Comments Acceped by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02677 2025-12-17 eess.IV cs.CV 79%

Multimodal Deep Learning for Stroke Prediction and Detection using Retinal Imaging and Clinical Data

基于视网膜成像和临床数据的多模态深度学习用于中风预测与检测

Saeed Shurrab, Aadim Nepal, Terrence J. Lee-St. John, Nicola G. Ghazi, Bartlomiej Piechowski-Jozwiak, Farah E. Shamout

机构 * Division of Engineering, New York University Abu Dhabi(纽约大学阿布扎赫尔分校工程系) Institute for Healthier Living Abu Dhabi(阿布扎赫尔健康生活研究所) Eye Institute at Cleveland Clinic Abu Dhabi(阿布扎赫尔克利夫兰医学中心眼科研究所) Canberra Hospital(堪培拉医院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出一种多模态深度学习方法,利用视网膜成像和临床数据预测中风风险,实验结果显示在中风检测和风险预测方面优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14489 2025-12-17 cs.CV 74%

SignIT: A Comprehensive Dataset and Multimodal Analysis for Italian Sign Language Recognition

SignIT:一个全面的数据集和多模态分析用于意大利手语识别

Alessia Micieli, Giovanni Maria Farinella, Francesco Ragusa

机构 * Computer Science - University of Catania, Italy(计算机科学 - 卡塔尼亚大学,意大利) Next Vision s.r.l., Spin-off of the University of Catania, Italy(2 Next Vision s.r.l.,卡塔尼亚大学分校,意大利)

专题命中 多模态评测 :multimodal(title);分类 cs.CV

AI总结 SignIT数据集通过多模态分析提升意大利手语识别的性能研究

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14496 2025-12-17 cs.CE 67%

BridgeNet: A Dataset of Graph-based Bridge Structural Models for Machine Learning Applications

BridgeNet: 一种基于图的桥梁结构模型数据集用于机器学习应用

Lazlo Bleker, Mustafa Cem Güneş, Pierluigi D'Acunto

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract)

AI总结 BridgeNet是一个包含20,000个桥梁结构模型的公开数据集,用于促进图机器学习和多模态学习在概念性结构设计中的应用。

Comments GNI Symposium on Artifical Intelligence for the Built World 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14574 2025-12-17 cs.CV cs.MM 62%

FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications

FoodLogAthl-218: 通过饮食管理应用构建一个现实世界的食物图像数据集

Mitsuki Watanabe, Sosuke Amano, Kiyoharu Aizawa, Yoko Yamakata

机构 * The University of Tokyo(东京大学) foo.log Inc.(foo.log公司)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.MM

AI总结 FoodLogAthl-218数据集通过用户真实餐食照片构建,包含218类食物图像及丰富元数据,提出基于上下文的分类任务和增量微调协议,用于提升饮食管理应用的图像分类性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13731 2025-12-17 cs.CV cs.AI 62%

Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

复杂数学表达式识别:基准测试、大规模数据集和强大基线

Weikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie, Zhineng Chen

机构 * \equalcontrib(机构1)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出CMER-Bench基准测试和大规模数据集MER-17M和CMER-3M,开发了CMERNet模型,有效提升了复杂数学表达式识别的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11827 2025-12-17 cs.CY cs.AI cs.CV 62%

Assessing Greenspace Attractiveness with ChatGPT, Claude, and Gemini: Do AI Models Reflect Human Perceptions?

利用ChatGPT、Claude和Gemini评估绿地吸引力:AI模型能反映人类感知吗?

Milad Malekzadeh, Magdalena Biernacka, Elias Willberg, Jussi Torkko, Edyta Łaszkiewicz, Tuuli Toivonen

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了AI模型在评估绿地吸引力方面的表现,发现其在正式绿地和非正式空间的判断一致性较高,但存在对安全性和本地嵌入质量的低估问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14683 2025-12-17 cs.LG 50%

Early Warning Index for Patient Deteriorations in Hospitals

医院患者恶化的早期预警指数

Dimitris Bertsimas, Yu Ma, Kimberly Villalobos Carballo, Gagan Singh, Michal Laskowski, Jeff Mather, Dan Kombert, Howard Haronian

机构 * Sloan School of Management, Massachusetts Institute of Technology(麻省理工学院斯隆管理学院) Operations and Information Management, University of Wisconsin Madison(威斯康星大学麦迪逊分校运营与信息管理系) Technology Management and Innovation, New York University(纽约大学技术管理与创新系) Hartford HealthCare(哈特福德医疗集团) Holistic Hospital Optimization(综合医院优化)

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出Early Warning Index模型,通过多模态机器学习预测患者恶化风险,结合SHAP解释提升可解释性,用于医院分诊和资源调度优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13903 2025-12-17 cs.RO 50%

PrediFlow: A Flow-Based Prediction-Refinement Framework for Real-Time Human Motion Prediction in Human-Robot Collaboration

PrediFlow: 一种基于流的预测-细化框架,用于人机协作中的实时人体运动预测

Sibo Tian, Minghui Zheng, Xiao Liang

机构 * J. Mike Walker ’66 Department of Mechanical Engineering, Texas A&M University(J. Mike Walker ’66机械工程系,德克萨斯农工大学) Zachry Department of Civil and Environmental Engineering, Texas A&M University(Zachry土木与环境工程系,德克萨斯农工大学)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 PrediFlow通过整合人类和机器人运动数据,提出了一种基于流的预测-细化框架,以提升实时人机协作中人体运动预测的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏