arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9176 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9176 篇

2411.02006 2025-09-16 cs.AI 79%

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Biao Wu, Yanda Li, Zhiwei Zhang, Yunchao Wei, Meng Fang, Ling Chen

机构 * Australian Artificial Intelligence Institute(澳大利亚人工智能研究所) The Pennsylvania State University(宾夕法尼亚州立大学) Beijing Jiaotong University(北京交通大学) University of Liverpool(利物浦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09683 2025-09-15 cs.IR cs.AI 79%

Forecasting Clicks in Digital Advertising: Multimodal Inputs and Interpretable Outputs

Briti Gangopadhyay, Zhao Wang, Shingo Takamatsu

机构 * Sony Group Corporation(索尼集团公司) Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎- Rocquencourt 国家信息与自动化研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13818 2025-09-15 cs.CV cs.LG 79%

Building Age Estimation: A New Multi-Modal Benchmark Dataset and Community Challenge

Nikolaos Dionelis, Alessandra Feliciotti, Mattia Marconcini, Devis Peressutti, Nika Oman Kadunc, JaeWan Park, Hagai Raja Sinulingga, Steve Andreas Immanuel, Ba Tran, Caroline Arnold, Nicolas Longépé

机构 * European Space Agency(欧洲航天局) Φ \Phi -lab(Phi实验室) ESRIN(ESRIN研究所) MindEarth(MindEarth公司) Sinergise/ Planet(Sinergise/Planet公司) TelePIX(TelePIX公司) Axelspace Corporation(Axelspace公司) Helmholtz Institute Hereon(海德堡研究所) German Climate Computing Center DKRZ(德国气候计算中心DKRZ)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 20 figures, 1 table, Submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09190 2025-09-12 cs.CV 79%

VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results

Hanwei Zhu, Haoning Wu, Zicheng Zhang, Lingyu Zhu, Yixuan Li, Peilin Chen, Shiqi Wang, Chris Wei Zhou, Linhan Cao, Wei Sun, Xiangyang Zhu, Weixia Zhang, Yucheng Zhu, Jing Liu, Dandan Zhu, Guangtao Zhai, Xiongkuo Min, Zhichao Zhang, Xinyue Li, Shubo Xu, Anh Dao, Yifan Li, Hongyuan Yu, Jiaojiao Yi, Yiding Tian, Yupeng Wu, Feiran Sun, Lijuan Liao, Song Jiang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ICCV VQualA Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08694 2025-09-11 cs.CV 79%

Multi-Modal Robust Enhancement for Coastal Water Segmentation: A Systematic HSV-Guided Framework

Zhen Tian, Christos Anagnostopoulos, Qiyuan Wang, Zhiwei Gao

机构 * School of Computing Science, University of Glasgow(格拉斯哥大学计算科学学院) School of Computing Engineering, University of Glasgow(格拉斯哥大学计算工程学院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08024 2025-09-11 cs.CV cs.CY 79%

Two Stage Context Learning with Large Language Models for Multimodal Stance Detection on Climate Change

Lata Pangtey, Omkar Kabde, Shahid Shafi Dar, Nagendra Kumar

机构 * Department of Computer Science and Engineering, Indian Institute of Technology (IIT) Indore(计算机科学与工程系,印度理工学院(IIT)印多尔) Chaitanya Bharathi Institute of Technology, Gandipet, Hyderabad, 500075(恰伊坦尼亚·巴哈提技术学院,加迪皮特,海得拉巴,500075)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06011 2025-09-10 cs.CV 79%

Light-Weight Cross-Modal Enhancement Method with Benchmark Construction for UAV-based Open-Vocabulary Object Detection

Zhenhai Weng, Xinjie Li, Can Wu, Weijie He, Jianfeng Lv, Dong Zhou, Zhongliang Yu

机构 * School of Automation, Chongqing University(重庆大学自动化学院) School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) Department of Control Science and Engineering, Harbin Institute of Technology(哈尔滨工业大学控制科学与工程系) Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong(香港中文大学机械与自动化工程系)

专题命中 多模态评测 :cross-modal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17471 2025-09-10 cs.CL 79%

FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain

Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao

机构 * Key Laboratory of High Confidence Software Technologies, CS, Peking University, China(北京大学高可信软件技术重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimedia Information Processing, School of Computer Sciences, Peking University(北京大学多媒体信息处理国家重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02100 2025-09-09 cs.HC cs.CL 79%

E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness

Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah

机构 * Centre for AI and ML, Edith Cowan University(人工智能与机器学习中心,埃德温·科温大学) University of Manchester(曼彻斯特大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments 15 pages, 4 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13470 2025-09-09 eess.SP cs.CV cs.LG 79%

Multimodal Latent Fusion of ECG Leads for Early Assessment of Pulmonary Hypertension

Mohammod N. I. Suvon, Shuo Zhou, Prasun C. Tripathi, Wenrui Fan, Samer Alabed, Bishesh Khanal, Venet Osmani, Andrew J. Swift, Chen, Chen, Haiping Lu

机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) Centre for Machine Intelligence, University of Sheffield(谢菲尔德大学智能中心) Department of Electrical & Computer Science Engineering, IITRAM Ahmedabad(印度阿赫迈德亚布德IITRAM电子与计算机科学工程系) Digital Environment Research Institute, Queen Mary University of London(伦敦玛丽女王大学数字环境研究院) Nepal Applied Mathematics and Informatics Institute for research (NAAMII), Nepal(尼泊尔应用数学与信息技术研究所) Department of Computing, Imperial College London(伦敦帝国学院计算机系) School of Medicine and Population Health, University of Sheffield(谢菲尔德大学医学与人口健康学院) Department of Clinical Radiology, Sheffield Teaching Hospitals(谢菲尔德教学医院放射科) National Institute for Health and Care Research (NIHR), Sheffield Biomedical Research Centre(英国国家健康与护理研究所(NIHR)谢菲尔德生物医学研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05773 2025-09-09 cs.CV 79%

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters

Zijian Chen, Wenjie Hua, Jinhao Li, Lirong Deng, Fan Du, Tingzhu Chen, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Lab(上海人工智能实验室) Wuhan University(武汉大学) East China Normal Unversity(华东师范大学) Macao Polytechnic University(澳门 polytechnic university) Southern University of Science and Technology(南方科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05330 2025-09-09 cs.AI 79%

MVRS: The Multimodal Virtual Reality Stimuli-based Emotion Recognition Dataset

Seyed Muhammad Hossein Mousavi, Atiye Ilanloo

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04823 2025-09-08 cs.SI cs.CL 79%

Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social Media

Yujie Wang, Yunwei Zhao, Jing Yang, Han Han, Shiguang Shan, Jie Zhang

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全国家重点实验室,计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19493 2025-09-04 cs.CR cs.CV 79%

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, Dongliang Xu

专题命中 多模态评测 :MLLM(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05779 2025-09-04 cs.CV 79%

A 3D Multimodal Feature for Infrastructure Anomaly Detection

Yixiong Jing, Wei Lin, Brian Sheil, Sinan Acikgoz

机构 * Department of Engineering Science, University of Oxford(工程科学系,牛津大学) Department of Geotechnical Engineering, College of Civil Engineering, Tongji University(地质工程系,同济大学土木学院) Construction Engineering, University of Cambridge(建设工程,剑桥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20356 2025-09-03 cs.CV 79%

Detecting Visual Information Manipulation Attacks in Augmented Reality: A Multimodal Semantic Reasoning Approach

Yanming Xiu, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments The paper has been accepted to the 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02256 2025-09-03 cs.CV 79%

A Multimodal Cross-View Model for Predicting Postoperative Neck Pain in Cervical Spondylosis Patients

Jingyang Shan, Qishuai Yu, Jiacen Liu, Shaolin Zhang, Wen Shen, Yanxiao Zhao, Tianyi Wang, Xiaolin Qin, Yiheng Yin

机构 * Chengdu Institute of Computer Applications, Chinese Academy of Sciences(成都计算机应用研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Department of Neurosurgery, the First Medical Center, Chinese PLA General Hospital(中国人民解放军总医院神经外科部) School of Medicine, Nankai University(南开大学医学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00687 2025-09-03 cs.CL 79%

Text Reinforcement for Multimodal Time Series Forecasting

Chen Su, Yuanhe Tian, Yan Song, Yongdong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) University of Washington(华盛顿大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17098 2025-09-03 cs.HC cs.AI 79%

CLARE: Cognitive Load Assessment in REaltime with Multimodal Data

Anubhav Bhatti, Prithila Angkan, Behnam Behinaein, Zunayed Mahmud, Dirk Rodenburg, Heather Braund, P. James Mclellan, Aaron Ruberto, Geoffery Harrison, Daryl Wilson, Adam Szulewski, Dan Howes, Ali Etemad, Paul Hungler

机构 * Department of Electrical and Computer Engineering and Ingenuity Research Labs(电气与计算机工程系和Ingenuity研究实验室) Queen’s University(女王大学) Ingenuity Research Labs(Ingenuity研究实验室) Department of Psychology(心理学系) School of Medicine(医学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 10 figures, 6 tables

Journal ref IEEE Transactions on Cognitive and Developmental Systems, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21793 2025-09-01 cs.LG cs.AI 79%

MoE-Health: A Mixture of Experts Framework for Robust Multimodal Healthcare Prediction

Xiaoyang Wang, Christopher C. Yang

机构 * Drexel University(德雷塞尔大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to The 16th ACM Conference on Bioinformatics, Computational Biology, and Health Informatics (ACM-BCB 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21169 2025-09-01 cs.CV 79%

SYNBUILD-3D: A large, multi-modal, and semantically rich synthetic dataset of 3D building models at Level of Detail 4

Kevin Mayer, Alex Vesel, Xinyi Zhao, Martin Fischer

机构 * Department of Civil and Environmental Engineering, Stanford University(土木与环境工程系,斯坦福大学) Department of Computer Science, Stanford University(计算机科学系,斯坦福大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14195 2025-08-29 cs.HC cs.CV 79%

A multimodal dataset for understanding the impact of mobile phones on remote online virtual education

Roberto Daza, Alvaro Becerra, Ruth Cobos, Julian Fierrez, Aythami Morales

机构 * Biometrics and Data Pattern Analytics Laboratory(生物信息与数据模式分析实验室) School of Engineering(工程学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Published in Scientific Data (Nature). GitHub repository of the dataset at: https://github.com/BiDAlab/IMPROVE

Journal ref Scientific Data (2025) 12:1332

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19862 2025-08-28 cs.CV cs.LG 79%

Multimodal Conditional MeshGAN for Personalized Aneurysm Growth Prediction

Long Chen, Ashiv Patel, Mengyun Qiao, Mohammad Yousuf Salmasi, Salah A. Hammouche, Vasilis Stavrinides, Jasleen Nagi, Soodeh Kalaie, Xiao Yun Xu, Wenjia Bai, Declan P. O'Regan

机构 * MRC Laboratory of Medical Sciences Imperial College London(医学科学实验室 Imperial College London) Imperial College Healthcare NHS Trust(帝国理工医疗 NHS Trust) Department of Mechanical Engineering University College London(机械工程系 University College London) London Postgraduate School of Surgery NHS England(伦敦外科研究生学校 NHS England) Faculty of Medicine Imperial College London(医学系 Imperial College London) Department of Chemical Engineering Imperial College London(化学工程系 Imperial College London) Department of Brain Sciences&Computing Imperial College London(脑科学与计算系 Imperial College London)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18608 2025-08-27 cs.AI 79%

eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases

Janet Wang, Xin Hu, Yunbei Zhang, Diabate Almamy, Vagamon Bamba, Konan Amos Sébastien Koffi, Yao Koffi Aubin, Zhengming Ding, Jihun Hamm, Rie R. Yotsu

机构 * Tulane University(Tulane大学) Bakke Graduate University(Bakke研究生大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18506 2025-08-27 cs.CV 79%

DoGFlow: Self-Supervised LiDAR Scene Flow via Cross-Modal Doppler Guidance

Ajinkya Khoche, Qingwen Zhang, Yixi Cai, Sina Sharif Mansouri, Patric Jensfelt

机构 * KTH Royal Institute of Technology(皇家理工学院) Autonomous Transport Solutions Lab, Scania Group(自主运输解决方案实验室,斯堪尼亚集团)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00152 2025-08-27 cs.CL 79%

Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data

Ekaterina Borisova, Fabio Barth, Nils Feldhus, Raia Abu Ahmad, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Sebastian Möller

机构 * Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI)(德国人工智能研究中心有限公司) Technische Universität Berlin(柏林技术大学) BIFOLD Deutsche Telekom(德国电信) Common Crawl Foundation(Common Crawl基金会) Humboldt-Universität zu Berlin(柏林洪堡大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments TRL@ACL 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18372 2025-08-27 cs.CV 79%

OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding

Hieu Nguyen, Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science(科学大学) University of Dayton(戴维森大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07155 2025-08-27 cs.CV 79%

Meta-Learned Modality-Weighted Knowledge Distillation for Robust Multi-Modal Learning with Missing Data

Hu Wang, Salma Hassan, Yuyuan Liu, Congbo Ma, Yuanhong Chen, Qing Li, Jiahui Geng, Bingjie Wang, Yu Tian, Yutong Xie, Jodie Avery, Louise Hull, Ian Reid, Mohammad Yaqub, Gustavo Carneiro

机构 * University of Adelaide, Australia(澳大利亚阿德莱德大学) New York University Abu Dhabi, UAE(阿联酋纽约大学阿布扎克分校) University of Surrey, UK(英国萨里大学) Boai hospital of Zhongshan, China(中国中山博爱医院) University of Central Florida, USA(美国佛罗里达大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17922 2025-08-26 cs.RO cs.CV 79%

Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model

Bokai Ji, Jie Gu, Xiaokang Ma, Chu Tang, Jingmin Chen, Guangxia Li

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00258 2025-08-26 cs.AI 79%

Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs

Qianqi Yan, Hongquan Li, Shan Jiang, Yang Zhao, Xinze Guan, Ching-Chen Kuo, Xin Eric Wang

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校) eBay

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏