arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-16 至 2025-09-16 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 14 篇

2509.11112 2025-09-16 cs.NI cs.AI cs.ET cs.IT cs.LG math.IT 83%

Multi-Modal Sensing Aided mmWave Beamforming for V2V Communications with Transformers

Muhammad Baqer Mollah, Honggang Wang, Hua Fang

机构 * University of Massachusetts Dartmouth(马萨诸塞大学达特茅斯分校) Yeshiva University(耶鲁大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments 6 Pages, Accepted to present at 2025 IEEE Global Communications Conference (GLOBECOM), Taipei, Taiwan

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06516 2025-09-16 cs.LG cs.AI 83%

QualityFM: a Multimodal Physiological Signal Foundation Model with Self-Distillation for Signal Quality Challenges in Critically Ill Patients

Zongheng Guo, Tao Chen, Manuela Ferrario

机构 * Department of Electronics, Information and Bioengineering, Politecnico di Milano(电子、信息与生物工程系,米兰理工学院) State Key Laboratory of Industrial Control Technology, Zhejiang University(工业控制技术国家重点实验室,浙江大学)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.AI

Comments 11 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07030 2025-09-16 cs.CL cs.AI cs.CV cs.IR cs.LG 82%

FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering

Amirhossein Abaskohi, Spandana Gella, Giuseppe Carenini, Issam H. Laradji

机构 * Department of Computer Science(计算机科学系) The University of British Columbia(不列颠哥伦比亚大学) ServiceNow Research(ServiceNow研究)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11620 2025-09-16 cs.CL cs.CY 79%

AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment

Kun Li, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02006 2025-09-16 cs.AI 79%

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Biao Wu, Yanda Li, Zhiwei Zhang, Yunchao Wei, Meng Fang, Ling Chen

机构 * Australian Artificial Intelligence Institute(澳大利亚人工智能研究所) The Pennsylvania State University(宾夕法尼亚州立大学) Beijing Jiaotong University(北京交通大学) University of Liverpool(利物浦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11335 2025-09-16 cs.LG cond-mat.mtrl-sci 78%

MatQnA: A Benchmark Dataset for Multi-modal Large Language Models in Materials Characterization and Analysis

Yonghao Weng, Liqiang Gao, Linwu Zhu, Jian Huang

机构 * Department of Materials Engineering(材料工程系) Zhejiang University(浙江大学) Department of Data Intelligence(数据智能系) Shiyanjia Lab of Scientific Compass(科学之桥实验室)

专题命中 多模态评测 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20358 2025-09-16 cs.LG 78%

Developing a Multi-Modal Machine Learning Model For Predicting Performance of Automotive Hood Frames

Abhishek Indupally, Satchit Ramnath

专题命中 多模态评测 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11136 2025-09-16 cs.LG cs.AI cs.CL cs.SI 76%

Agentic Username Suggestion and Multimodal Gender Detection in Online Platforms: Introducing the PNGT-26K Dataset

Farbod Bijary, Mohsen Ebadpour, Amirhosein Tajbakhsh

机构 * Amirkabir University of Technology(阿米尔卡比尔理工大学) Iran University of Science & Technology(伊朗科学技术大学)

专题命中 多模态评测 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10683 2025-09-16 cs.CV cs.AI 62%

A Comparison and Evaluation of Fine-tuned Convolutional Neural Networks to Large Language Models for Image Classification and Segmentation of Brain Tumors on MRI

Felicia Liu, Jay J. Yoo, Farzad Khalvati

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11589 2025-09-16 cs.CV 57%

MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment

Yanyun Pu, Kehan Li, Zeyi Huang, Zhijie Zhong, Kaixiang Yang

机构 * Huawei Technologies Co.(华为技术有限公司) South China University of Technology(南方科技大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11459 2025-09-16 cs.AI 57%

Knowledge-Guided Adaptive Mixture of Experts for Precipitation Prediction

Chen Jiang, Kofi Osei, Sai Deepthi Yeddula, Dongji Feng, Wei-Shinn Ku

机构 * Samuel Ginn College of Engineering(萨姆uel吉恩工程学院) Auburn University(阿伯杜大学) Gustavus Adolphus College(古斯塔夫·阿道夫学院) School of Engineering and Computer Science(工程与计算机科学学院) Oakland University(奥克兰大学) School of Computing and Design(计算与设计学院) California State University Monterey Bay(蒙特利尔湾加州州立大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10570 2025-09-16 cs.RO cs.AI 57%

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

Wei Dai, Shengen Wu, Wei Wu, Zhenhao Wang, Sisuo Lyu, Haicheng Liao, Limin Yu, Weiping Ding, Runwei Guan, Yutao Yue

机构 * Department of Mathematical Sciences, School of Physical sciences, University of Liverpool(利物浦大学数学科学系) Department of Communications and Networking, School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学通讯与网络系) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) Deep Interdisciplinary Intelligence Lab, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)深度跨学科智能实验室) School of Mathematics and Statistics, Shandong University(山东大学数学与统计学院) School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)数据科学与分析方向) Institute of Deep Perception Technology, Jiangsu(江苏深度感知技术研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11876 2025-09-16 cs.HC 50%

Lost in Data: How Older Adults Perceive and Navigate Health Data Representations

Peterson Jean, Emma Murphy, Enda Bates

专题命中 多模态评测 :multimodal(abstract)

Comments AAATE 2025 Proceedings (Research Strand). Licensed under CC BY-NC-ND 4.0. ISBN: 978-9925-604-07-4

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10556 2025-09-16 q-bio.TO cs.CE 50%

COVID-BLUeS -- A Prospective Study on the Value of AI in Lung Ultrasound Analysis

Nina Wiedemann, Dianne de Korte-de Boer, Matthias Richter, Sjors van de Weijer, Charlotte Buhre, Franz A. M. Eggert, Sophie Aarnoudse, Lotte Grevendonk, Steffen Röber, Carlijn M. E. Remie, Wolfgang Buhre, Ronald Henry, Jannis Born

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏