arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-24 至 2025-10-24 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2501.01243 2025-10-24 cs.CV cs.AI cs.CL 82%

Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants

Lixiong Qin, Shilong Ou, Miaoxuan Zhang, Jiangning Wei, Yuhang Zhang, Xiaoshuai Song, Yuchen Liu, Mei Wang, Weiran Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 50 pages, 14 figures, 42 tables. NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20381 2025-10-24 cs.CL cs.AI 81%

VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation

Son T. Luu, Trung Vo, Hiep Nguyen, Khanh Quoc Tran, Kiet Van Nguyen, Vu Tran, Ngan Luu-Thuy Nguyen, Le-Minh Nguyen

机构 * Japan Advanced Institute of Science and Technology(日本先进科学研究院) University of Information Technology(信息技术大学) Vietnam National University(越南国家大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments VLSP 2025 MLQA-TSR Share Task

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19892 2025-10-24 cs.CL cs.AI 81%

Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities

Nishant Balepur, Dang Nguyen, Dayeon Ki

机构 * University of Maryland(马里兰大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted as a Spotlight paper at the EMNLP 2025 Wordplay Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20632 2025-10-24 cs.AI 79%

Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications

Shuyi Xie, Ziqin Liew, Hailing Zhang, Haibo Zhang, Ling Hu, Zhiqiang Zhou, Shuman Liu, Anxiang Zeng

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11520 2025-10-24 cs.CV 79%

mmWalk: Towards Multi-modal Multi-view Walking Assistance

Kedi Ying, Ruiping Liu, Chongyan Chen, Mingzhe Tao, Hao Shi, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

机构 * CV:HCI, KIT(KIT计算机视觉与人机交互中心) Hunan University(湖南大学) ETH Zurich(苏黎世联邦理工学院) University of Texas at Austin(德克萨斯大学奥斯汀分校) Zhejiang University(浙江大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track. Data and Code: https://github.com/KediYing/mmWalk

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04462 2025-10-24 cs.CL cs.AI 62%

Benchmarking GPT-5 for biomedical natural language processing

Yu Hou, Zaifu Zhan, Min Zeng, Yifan Wu, Shuang Zhou, Rui Zhang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20612 2025-10-24 cs.CV cs.CL cs.LG 62%

Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models

Peter Robicheaux, Matvei Popov, Anish Madan, Isaac Robinson, Joseph Nelson, Deva Ramanan, Neehar Peri

机构 * Roboflow Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments The first two authors contributed equally. This work has been accepted to the Neural Information Processing Systems (NeurIPS) 2025 Datasets & Benchmark Track. Project Page: https://rf100-vl.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16793 2025-10-24 cs.CV 57%

REOBench: Benchmarking Robustness of Earth Observation Foundation Models

Xiang Li, Yong Tao, Siyuan Zhang, Siwei Liu, Zhitong Xiong, Chunbo Luo, Lu Liu, Mykola Pechenizkiy, Xiao Xiang Zhu, Tianjin Huang

机构 * University of Bristol, UK(英国布里斯托大学) University of Exeter, UK(英国埃克塞特大学) South China Normal University, China(华南师范大学) The University of Aberdeen, UK(英国阿伯丁大学) Technical University of Munich, Germany(慕尼黑技术大学) Eindhoven University of Technology, NL(埃因霍温理工大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeruIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15560 2025-10-24 cs.CV cs.HC eess.IV eess.SP 57%

QUB-PHEO: A Visual-Based Dyadic Multi-View Dataset for Intention Inference in Collaborative Assembly

Samuel Adebayo, Seán McLoone, Joost C. Dessing

机构 * Centre for Intelligent Autonomous Manufacturing Systems, Queen’s University Belfast(智能自主制造系统研究中心,女王大学贝尔法斯特) School of Electronics, Electrical Engineering and Computer Science, Queen’s University Belfast(电子、电气工程与计算机科学学院,女王大学贝尔法斯特) School of Psychology, Queen’s University Belfast(心理学学院,女王大学贝尔法斯特)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Journal ref IEEE Access, Vol. 12, pp. 157050-157066, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06259 2025-10-24 cs.CY cs.LG 50%

Beyond Static Knowledge Messengers: Towards Adaptive, Fair, and Scalable Federated Learning for Medical AI

Jahidul Arafat, Fariha Tasmin, Sanjaya Poudel, Iftekhar Haider

机构 * Department of Computer Science and Software Engineering, Auburn University(计算机科学与软件工程系,阿伯丁大学) Department of Information and Communication Technology, Bangladesh University of Professionals(信息与通信技术系,孟加拉国专业大学) Mymensingh Medical College and Hospital(迈明辛医疗学院和医院)

专题命中 多模态评测 :multi-modal(abstract)

Comments 20 pages, 4 figures, 14 tables. Proposes Adaptive Fair Federated Learning (AFFL) algorithm and MedFedBench benchmark suite for healthcare federated learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12816 2025-10-24 quant-ph 50%

Floquet Analysis of Frequency Collisions

Kentaro Heya, Moein Malekakhlagh, Seth Merkel, Naoki Kanazawa, Emily Pritchett

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏