arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-27 至 2025-11-27 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9 篇

2501.12422 2025-11-27 cs.LG cs.AI cs.CV 90%

CroMe: Multimodal Fake News Detection using Cross-Modal Tri-Transformer and Metric Learning

CroMe:基于跨模态三变换器和度量学习的多模态假新闻检测

Eunjee Choi, Junhyun Ahn, XinYu Piao, Jong-Kook Kim

机构 * Department of Electrical and Computer Engineering, Korea University, Seoul 02841, Republic of Korea(电气与计算机工程系,韩国大学,首尔)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 CroMe通过跨模态三变换器和度量学习方法,提升多模态假新闻检测的准确性和有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21339 2025-11-27 cs.CV cs.AI 81%

SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding

SurgMLLMBench: 一个用于手术场景理解的多模态大语言模型基准数据集

Tae-Min Choi, Tae Kyeong Jeong, Garam Kim, Jaemin Lee, Yeongyoon Koh, In Cheul Choi, Jae-Ho Chung, Jong Woong Park, Juyoun Park

机构 * Samsung Research(三星研究所) Center for Humanoid Research, Korea Institute of Science and Technology(人类研究学院,韩国科学技术院) Department of plastic surgery, College of medicine, Korea University(医学院整形外科部,韩国大学) Department of orthopedic surgery, College of medicine, Korea University(医学院骨科部,韩国大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 SurgMLLMBench是一个用于手术场景理解的多模态大语言模型基准数据集,整合了像素级分割和结构化VQA注释,支持跨领域的全面评估和更丰富的视觉-对话交互。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12307 2025-11-27 cs.CV cs.CL 81%

LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?

LogicOCR: 大型多模态模型在文本丰富的图像上逻辑推理是否表现优异?

Maoyuan Ye, Haibin He, Qihuang Zhong, Jing Zhang, Juhua Liu, Bo Du

机构 * School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(计算机学院、多媒体软件国家工程研究中心、人工智能研究院、多媒体与网络通信工程湖北省重点实验室、武汉大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 LogicOCR通过构建包含生成和现实图像问题的基准测试,评估大型多模态模型在文本丰富图像上的逻辑推理能力,并提出TextCue方法提升模型对关键文本区域的感知。

Comments GitHub: https://github.com/MiliLab/LogicOCR

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21364 2025-11-27 cs.LG cs.CV 80%

BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla

BanglaMM-Disaster: 一种基于Transformer的多模态深度学习框架,用于孟加拉语多类灾害分类

Ariful Islam, Md Rifat Hossen, Md. Mahmudul Arif, Abdullah Al Noman, Md Arifur Rahman

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Chittagong University of Engineering and Technology(奇特格隆工程与技术大学) Department of Electronics and Telecommunication Engineering(电子与电信工程系) Wilmington University(维明顿大学) College of Graduate and Professional Studies(研究生与专业研究学院) Trine University(特林大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 BanglaMM-Disaster通过结合文本和视觉数据,提出了一种多模态深度学习框架,用于孟加拉语多类灾害分类,提升了灾害响应效率。

Comments Presented at the 2025 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON), November 21-22, 2025, University of Rajshahi, Bangladesh. 6 pages, 9 disaster classes, multimodal dataset with 5,037 samples

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21422 2025-11-27 cs.CV 79%

E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework

E-M3RF:一个等价多模态3D重装配框架

Adeela Islam, Stefano Fiorini, Manuel Lecha, Theodore Tsesmelis, Stuart James, Pietro Morerio, Alessio Del Bue

机构 * Fondazione Istituto Italiano di Tecnologia(意大利技术研究院) University of Genova(热那亚大学) Durham University(杜伦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 E-M3RF通过多模态特征融合和SE(3)流匹配,有效提升3D碎片重装配的精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20701 2025-11-27 cs.AI cs.LG 79%

Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework

跨领域评估不同数据集的多模态链式推理在Amazon CoT框架中的表现

Nitya Tiwari, Parv Maheshwari, Vidisha Agarwal

机构 * Indian Institute of Technology Bombay(印度理工学院孟买学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文评估了多模态链式推理在不同数据集上的表现,发现视觉整合减少幻觉但不同问题类型对推理有效性影响显著,为多模态推理系统改进提供参考。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15277 2025-11-27 cs.CL 70%

Web-Shepherd: Advancing PRMs for Reinforcing Web Agents

Web-Shepherd: 促进强化网络代理的PRMs

Hyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim, Seungjun Moon, Gyeom Hwangbo, Dongha Lim, Minjin Kim, Yeonjun Hwang, Minju Gwak, Dongwook Choi, Minseok Kang, Gwanhoon Im, ByeongUng Cho, Hyojun Kim, Jun Hee Han, Taeyoon Kwon, Minju Kim, Beong-woo Kwak, Dongjin Kang, Jinyoung Yeo

机构 * Georgia Institute of Technology(佐治亚理工学院) Department of Artificial Intelligence, Yonsei University(延世大学人工智能系) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL

AI总结 Web-Shepherd是首个用于强化网络代理的PRM,通过构建大规模数据集和元评估基准,提升了网络导航任务的准确性和效率。

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14914 2025-11-27 cs.CV cs.CL cs.LG 62%

CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness

CAPability: 一个全面的视觉描述基准,用于评估正确性和全面性

Zhihang Liu, Chen-Wei Xie, Bin Wen, Feiwu Yu, Jixuan Chen, Pandeng Li, Boqiang Zhang, Nianzu Yang, Yinglu Li, Zuan Gao, Yun Zheng, Hongtao Xie

机构 * University of Science and Technology of China(中国科学技术大学) Tongyi Lab, Alibaba Group(阿里云实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 CAPability是一个全面的视觉描述基准,通过12个维度评估描述的正确性和全面性,利用人类标注数据和启发式指标提升多模态模型的描述能力分析。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21444 2025-11-27 cs.AI physics.ao-ph 57%

EWE: An Agentic Framework for Extreme Weather Analysis

EWE:极端天气分析的代理框架

Zhe Jiang, Jiong Wang, Xiaoyu Yue, Zijie Guo, Wenlong Zhang, Fenghua Ling, Wanli Ouyang, Lei Bai

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The University of Sydney(悉尼大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 EWE是首个用于极端天气分析的智能代理框架,通过知识引导的规划和闭环推理实现自动化诊断,提供首个该领域基准测试,推动科学发现民主化进程。

详情

展开后加载摘要…

URL PDF HTML 收藏