arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-12 至 2025-12-12 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 7 篇

2506.11375 2025-12-12 cs.AI cs.CL 84%

Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables

在化学表格上评估多模态大语言模型的基准测试

Yitong Zhou, Mingyue Cheng, Qingyang Mao, Yucong Luo, Qi Liu, Yupeng Li, Xiaohan Zhang, Deguang Liu, Xin Li, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) University of Science and Technology of China(中国科学技术大学) Artificial Intelligence Research Institute(人工智能研究院) iFLYTEK Co., Ltd(iFLYTEK公司)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出ChemTable基准,用于评估多模态模型在理解化学表格中的能力,揭示了现有模型在跨模态对齐和领域推理方面的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10701 2025-12-12 cs.LG 82%

HybridVFL: Disentangled Feature Learning for Edge-Enabled Vertical Federated Multimodal Classification

HybridVFL:面向边缘计算的垂直联邦多模态分类的解耦特征学习

Mostafa Anoosha, Zeinab Dehghani, Kuniko Paxton, Koorosh Aslansefat, Dhavalkumar Thakker

机构 * University of Hull(赫尔大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

AI总结 HybridVFL通过客户端特征解耦与服务器跨模态Transformer融合,提升边缘计算中多模态分类的隐私保护性能。

Comments 6 pages, 2 figures, 1 table. Accepted at UCC '25 (IEEE/ACM 18th International Conference on Utility and Cloud Computing), December 1-4, 2025, Nantes, France. DOI to be activated upon final publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10791 2025-12-12 cs.CL cs.AI 62%

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

FACTS排行榜:大型语言模型事实性的综合基准

Aileen Cheng, Alon Jacovi, Amir Globerson, Ben Golan, Charles Kwong, Chris Alberti, Connie Tao, Eyal Ben-David, Gaurav Singh Tomar, Lukas Haas, Yonatan Bitton, Adam Bloniarz, Aijun Bai, Andrew Wang, Anfal Siddiqui, Arturo Bajuelos Castillo, Aviel Atias, Chang Liu, Corey Fry, Daniel Balle, Deepanway Ghosal, Doron Kukliansky, Dror Marcus, Elena Gribovskaya, Eran Ofek, Honglei Zhuang, Itay Laish, Jan Ackermann, Lily Wang, Meg Risdal, Megan Barnes, Michael Fink, Mohamed Amin, Moran Ambar, Natan Potikha, Nikita Gupta, Nitzan Katz, Noam Velan, Ofir Roval, Ori Ram, Polina Zablotskaia, Prathamesh Bang, Priyanka Agrawal, Rakesh Ghiya, Sanjay Ganapathy, Simon Baumgartner, Sofia Erell, Sushant Prakash, Thibault Sellam, Vikram Rao, Xuanhui Wang, Yaroslav Akulov, Yulong Yang, Zhen Yang, Zhixin Lai, Zhongru Wu, Anca Dragan, Avinatan Hassidim, Fernando Pereira, Slav Petrov, Srinivasan Venkatachary, Tulsee Doshi, Yossi Matias, Sasha Goldshtein, Dipanjan Das

机构 * Google(谷歌)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 FACTS排行榜为大型语言模型提供事实性评估的综合基准,通过四个子排行榜全面衡量模型在不同场景下的事实准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10336 2025-12-12 cs.CL cs.AI 62%

Multilingual VLM Training: Adapting an English-Trained VLM to French

多语言视觉语言模型训练:将英语训练的VLM适应到法语

Jules Lahmi, Alexis Roger

机构 * Ecole Polytechnique(巴黎政治学院) Mila - Quebec AI Institute(魁北克人工智能研究所) McGill University(麦吉尔大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了将英语训练的视觉语言模型适应到法语等其他语言的挑战,通过比较不同方法的性能和计算成本,发现数据翻译是多语言VLM性能的主要瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10592 2025-12-12 cs.CV 57%

Salient Object Detection in Complex Weather Conditions via Noise Indicators

通过噪声指示器进行复杂天气条件下的显著物体检测

Quan Chen, Xiaokai Yang, Tingyu Wang, Rongfeng Lu, Xichun Sheng, Yaoqi Sun, Chenggang Yan

机构 * School of Automation, Hangzhou Dianzi University(自动化学院,杭州电子大学) College of Artificial Intelligence, Jiaxing University(人工智能学院,嘉兴大学) School of Communication Engineering, Hangzhou Dianzi University(通信工程学院,杭州电子大学) School of Faculty of Applied Science, Macao Polytechnic University(应用科学学院,澳门理工学院) School of Mathematics and Computer Science, Lishui University(数学与计算机科学学院,丽水大学) Lishui Institute of Hangzhou Dianzi University(杭州电子大学丽水学院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种针对复杂天气条件的显著物体检测框架,通过引入噪声指示器融合模块提升分割准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10262 2025-12-12 cs.CV 57%

VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models

基于视觉大语言模型的新型类别发现

Yuetong Su, Baoguo Wei, Xinyu Wang, Xu Li, Lixin Li

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 基于视觉大语言模型的新型类别发现方法,通过融合视觉-文本语义和原型引导聚类,提升未知类别分类准确率25.3%,并展现对长尾分布的鲁棒性。

Comments 8 pages, 5 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21336 2025-12-12 cs.CV 57%

UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation

UniBiomed: 一种用于 grounded 生物医学图像解释的通用基础模型

Linshan Wu, Yuxiang Nie, Sunan He, Jiaxin Zhuang, Luyang Luo, Tao Li, Zhuoyao Xie, Dexuan Chen, Yinghua Zhao, Neeraj Mahboobani, Varut Vardhanabhuti, Ronald Cheong Kin Chan, Yifan Peng, Pranav Rajpurkar, Hao Chen

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 UniBiomed是一种基于多模态大语言模型和Segment Anything Model的通用基础模型,能够同时生成诊断结果并分割生物医学目标,提升生物医学图像分析的可解释性与性能。

Comments A universal foundation model for grounded biomedical image interpretation

详情

展开后加载摘要…

URL PDF HTML 收藏