arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-16 至 2025-12-16 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 16 篇

2503.06940 2025-12-16 cs.CV 85%

CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing

CineBrain: 一种大规模多模态大脑数据集用于自然视听叙事处理

Jianxiong Gao, Yichang Liu, Baofeng Yang, Jianfeng Feng, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 CineBrain通过多模态融合编码器和神经潜在解码器实现动态视频重建,利用fMRI和EEG互补优势提升视觉保真度,并揭示听觉输入对视觉感知的影响。

Comments Project Page: https://jianxgao.github.io/CineBrain

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10576 2025-12-16 cs.CV 83%

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs

HumanSense: 从多模态感知到通过推理的 empathetic 上下文感知响应

Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Yi Yuan, Jingdong Chen, Le Wang

专题命中 多模态评测 :multimodal(title,abstract);omni-modal(abstract);分类 cs.CV

AI总结 HumanSense通过多阶段模态渐进强化学习提升MLLMs在多模态感知和交互任务中的性能,强调推理能力对共情反馈的重要性。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13297 2025-12-16 cs.AI cs.LG 79%

MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data

MedInsightBench: 通过多步骤洞察发现评估医疗分析代理在多模态医疗数据中的能力

Zhenghao Zhu, Chuxue Cao, Sirui Han, Yuanfeng Song, Xing Chen, Caleb Chen Cao, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ByteDance(字节跳动)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.AI

AI总结 MedInsightBench通过多步骤洞察发现评估医疗分析代理在多模态医疗数据中的能力,提出MedInsightAgent框架提升医疗数据洞察发现性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12503 2025-12-16 cs.AI 57%

KidsArtBench: Multi-Dimensional Children's Art Evaluation with Attribute-Aware MLLMs

KidsArtBench: 多维度儿童艺术评估与属性感知的大语言模型

Mingrui Ye, Chanjin Zheng, Zengyi Yu, Chenyu Xiang, Zhixue Zhao, Zheng Yuan, Helen Yannakoudakis

机构 * King’s College London(伦敦国王学院) East China Normal University(东华大学) University of Sheffield(谢菲尔德大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 KidsArtBench通过属性感知的多LoRA方法提升儿童艺术评估的准确性与教育意义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12500 2025-12-16 cs.HC cs.AI 57%

Explainable AI as a Double-Edged Sword in Dermatology: The Impact on Clinicians versus The Public

可解释AI在皮肤科中的双刃剑作用:对医生与公众的影响

Xuhai Xu, Haoyu Hu, Haoran Zhang, Will Ke Wang, Reina Wang, Luis R. Soenksen, Omar Badri, Sheharbano Jafry, Elise Burger, Lotanna Nwandu, Apoorva Mehta, Erik P. Duhaime, Asif Qasim, Hause Lin, Janis Pereira, Jonathan Hershon, Paulius Mui, Alejandro A. Gru, Noémie Elhadad, Lena Mamykina, Matthew Groh, Philipp Tschandl, Roxana Daneshjou, Marzyeh Ghassemi

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 研究探讨了可解释AI在皮肤科诊断中的影响,发现不同专业背景的人对XAI的反应不同,LLM在医疗AI中具有双刃剑效应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12444 2025-12-16 cs.CL 57%

Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors

GPT能否替代人类评分者?机器生成的隐喻规范的有效性和可靠性

Veronica Mangiaterra, Hamad Al-Azary, Chiara Barattieri di San Pietro, Paolo Canal, Valentina Bambini

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 本文研究了GPT在隐喻评分中的有效性与可靠性,发现较大模型能有效替代人类评分,但需注意隐喻惯例性和多模态因素的影响。

Comments 30 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12443 2025-12-16 cs.AI cs.SE 57%

AI Transparency Atlas: Framework, Scoring, and Real-Time Model Card Evaluation Pipeline

AI透明度地图:框架、评分与实时模型卡片评估流程

Akhmadillo Mamirov, Faiaz Azmain, Hanyu Wang

机构 * Department of Computer Science, The College of Wooster, Wooster, OH, USA(计算机科学系,沃斯特学院,沃斯特,OH,USA) Robert F. Wagner Graduate School of Public Service, New York University, New York, NY, USA(罗伯特·F·瓦格纳公共事务研究生学院,纽约大学,纽约,NY,USA)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 本文提出AI透明度框架和实时评估流程,评估50个模型发现前沿实验室合规率约80%,而多数供应商低于60%,安全关键领域存在显著缺口。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12424 2025-12-16 cs.CV cs.LG 57%

ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics

ViInfographicVQA:单图和多图越南信息图视觉问答基准测试

Tue-Thu Van-Dinh, Hoang-Duy Tran, Truong-Binh Duong, Mai-Hanh Pham, Binh-Nam Le-Nguyen, Quoc-Thai Nguyen

机构 * Neu.edu.vn

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 ViInfographicVQA是首个针对越南信息图VQA的基准测试,通过单图和多图任务评估模型在复杂信息图中的视觉问答能力。

Comments 10 pages, 4 figures, Accepted to AI4Research @ AAAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04457 2025-12-16 cs.CL 57%

Do MLLMs Really Understand the Charts?

多模态大语言模型真的能理解图表吗?

Xiao Zhang, Dongyuan Li, Liuyu Xiang, Yao Zhang, Cheng Zhong, Zhaofeng He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) AAITC, CTO Organization, Lenovo(AAITC、CTO组织、联想) Lenovo(联想)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 本文提出 ChartVR-3B/7B 模型,通过 VR-RFT 策略提升图表视觉推理能力,并在 ChartVRBench 上取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13405 2025-12-16 cs.CL 57%

RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis

RealHiTBench: 一个全面的现实感分层表格基准,用于评估基于LLM的表格分析

Pengzuo Wu, Yuhang Yang, Guangcheng Zhu, Chao Ye, Hong Gu, Xu Lu, Ruixuan Xiao, Bowen Bao, Yijing He, Liangyu Zha, Wentao Ye, Junbo Zhao, Haobo Wang

机构 * Zhejiang University(浙江大学) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司) Institute of Computing Innovation, Zhejiang University(浙江大学计算机创新研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 RealHiTBench是一个用于评估基于LLM的表格分析能力的综合现实感分层表格基准,通过多种输入格式和复杂结构的表格测试,验证了改进LLMs对表格层次结构感知的重要性。

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17159 2025-12-16 cs.CV 57%

RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness

RobustMerge: 用于MLLMs的参数高效模型合并的鲁棒性

Fanhu Zeng, Haiyang Guo, Fei Zhu, Li Shen, Hao Tang

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Centre for Artificial Intelligence and Robotics, HKISI-CAS(香港科学院智能机器人中心) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Shenzhen Loop Area Institute(深圳河套学院) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理国家重点实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 RobustMerge提出一种无需训练的参数高效模型合并方法,通过参数修剪和跨任务归一化提升多任务泛化能力。

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12059 2025-12-16 cs.AI 57%

The Forecast Critic: Leveraging Large Language Models for Poor Forecast Identification

预测批评者:利用大语言模型进行差预测识别

Luke Bhan, Hanyu Zhang, Andrew Gordon Wilson, Michael W. Mahoney, Chuck Arvin

机构 * Amazon(亚马逊)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

AI总结 利用大语言模型进行预测监控,通过识别不合理预测提升预测质量。

Comments Presented at AAAI 2026 AI4TS workshop and AABA4ET workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17699 2025-12-16 cs.CV 57%

Establishing Reality-Virtuality Interconnections in Urban Digital Twins for Superior Intelligent Road Inspection and Simulation

在城市数字孪生中建立现实-虚拟联系以实现更智能的道路检查与模拟

Yikang Zhang, Chuang-Wei Liu, Jiahang Li, Yingbing Chen, Jie Cheng, Rui Fan

机构 * College of Electronics & Information Engineering, Shanghai Research Institute for Intelligent Autonomous Systems, the State Key Laboratory of Intelligent Autonomous Systems, and Frontiers Science Center for Intelligent Autonomous Systems, Tongji University, Shanghai 201804, China(电子与信息工程学院,上海智能自主系统研究院,智能自主系统国家重点实验室,智能自主系统前沿科学中心,同济大学,上海)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种多模态传感器平台与城市数字孪生结合的智能道路检查系统,通过高保真的道路缺陷场景提升驾驶任务的感知和决策性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13010 2025-12-16 cs.LG q-bio.TO 50%

Deep Learning-Driven Inversion Framework for Shear Modulus Estimation in Magnetic Resonance Elastography (DIME)

基于深度学习的磁共振弹性成像中剪切模量估计反演框架(DIME)

Hassan Iftikhar, Rizwan Ahmad, Arunark Kolipaka

机构 * Biomedical Engineering, The Ohio State University(生物医学工程,俄亥俄州立大学) Department of Radiology, The Ohio State University(放射学系,俄亥俄州立大学) Davis Heart & Lung Research Institute, The Ohio State University(戴维斯心脏与肺研究所,俄亥俄州立大学)

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出DIME框架,利用深度学习提升MRE中剪切模量估计的鲁棒性和精度,实验显示其在刚度图生成和真实值匹配方面优于传统MMDI方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15497 2025-12-16 eess.SP 50%

A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems

机械系统中空化强度识别的机器学习综述

Yu Sha, Ningtao Liu, Haofeng Liu, Junqi Tao, Zhenxing Niu, Guojun Huang, Yao Yao, Jiaqi Liang, Moxian Qian, Horst Stoecker, Domagoj Vnucec, Andreas Widl, Kai Zhou

专题命中 多模态评测 :multi-modal(abstract)

AI总结 本文综述了机械系统中空化强度识别的机器学习发展,强调传统方法与深度学习的演变,并展望未来在多源数据处理和工业应用中的发展方向。

Comments 43 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15562 2025-12-16 cs.RO 50%

Transformer-XL for Long Sequence Tasks in Robotic Learning from Demonstration

Transformer-XL在机器人学习从示范中的长序列任务应用

Gao Tianci

机构 * Bauman Moscow State Technical University(巴甫洛夫莫斯科国立技术大学)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 本文提出利用Transformer-XL处理机器人学习从示范中的长序列任务,通过多模态传感器输入提升任务成功率和计算效率。

Comments 11 pages,4 figures

Journal ref Tianci G. Transformer-xl for long sequence tasks in robotic learning from demonstrations[C]//International Conference on Intelligent Systems. Cham: Springer Nature Switzerland, 2024: 25-36

详情

展开后加载摘要…

URL PDF HTML 收藏