arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-28 至 2025-10-28 共收录 129 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 32 篇

2510.23451 2025-10-28 cs.CL cs.AI cs.CV 85%

Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences

Zhuoran Jin, Hongbang Yuan, Kejian Zhu, Jiachun Li, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态评测 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 48 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22443 2025-10-28 cs.CV cs.LG 83%

Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents

Vijay Veerabadran, Fanyi Xiao, Nitin Kamra, Pedro Matias, Joy Chen, Caley Drooff, Brett D Roads, Riley Williams, Ethan Henderson, Xuanyi Zhao, Kevin Carlberg, Joseph Tighe, Karl Ridgeway

机构 * Meta Reality Labs(Meta 现实实验室) Meta FAIR

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted as a spotlight paper at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12520 2025-10-28 cs.CV 83%

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan, Shuliang Liu, PeiJun Wu, Peijie Jiang, Jia Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) Ant Group, Alibaba(蚂蚁集团)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22527 2025-10-28 astro-ph.IM astro-ph.GA cs.LG 82%

Multi-Modal Masked Autoencoders for Learning Image-Spectrum Associations for Galaxy Evolution and Cosmology

Morgan Himes, Samiksha Krishnamurthy, Andrew Lizarraga, Srinath Saikrishnan, Vikram Seenivasan, Jonathan Soriano, Ying Nian Wu, Tuan Do

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract)

Comments 8 pages, 3 figures, 1 table, accepted to NeurIPS 2025 Workshop ML4PS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22622 2025-10-28 cs.CR cs.CV cs.MM 81%

DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection

Kangran Zhao, Yupeng Chen, Xiaoyu Zhang, Yize Chen, Weinan Guan, Baicheng Chen, Chengzhe Sun, Soumyya Kanti Datta, Qingshan Liu, Siwei Lyu, Baoyuan Wu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University at Buffalo, State University of New York(纽约州立大学布法罗分校) Nanjing University of Posts and Telecommunications(南京邮电大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20759 2025-10-28 cs.CV cs.AI 81%

PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding

Ansel Blume, Jeonghwan Kim, Hyeonjeong Ha, Elen Chatikyan, Xiaomeng Jin, Khanh Duy Nguyen, Nanyun Peng, Kai-Wei Chang, Derek Hoiem, Heng Ji

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California Los Angeles(加州大学洛杉矶分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 Spotlight; project page: https://wjdghks950.github.io/partonomy.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21076 2025-10-28 cs.CV 79%

DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding

Weihao Xuan, Junjue Wang, Heli Qi, Zihang Chen, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya

机构 * The University of Tokyo(东京大学) RIKEN AIP(理化学研究所AIP) Waseda University(早稻田大学) Wuhan University(武汉大学) Stanford University(斯坦福大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09252 2025-10-28 cs.CV 79%

Zero-Shot Multi-modal Large Language Model v.s. Supervised Deep Learning: A Comparative Study on CT-Based Intracranial Hemorrhage Subtyping

Yinuo Wang, Yue Zeng, Kai Chen, Cai Meng, Chao Pan, Zhouping Tang

机构 * Image Processing Center, Beihang University(北京航空航天大学图像处理中心) School of Mechanical Engineering and Automation, Beihang University(北京航空航天大学机械工程与自动化学院) Department of Neurology, Tongji Hospital, Tongji Medical College, Huazhong University of Science and Technology(华中科技大学同济医学院神经内科,同济医院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07246 2025-10-28 cs.LG cs.CV 79%

ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area

Junxian Li, Di Zhang, Xunzhi Wang, Zeying Hao, Jingdi Lei, Qian Tan, Cai Zhou, Wei Liu, Yaotian Yang, Xinrui Xiong, Weiyun Wang, Zhe Chen, Wenhai Wang, Wei Li, Shufei Zhang, Mao Su, Wanli Ouyang, Yuqiang Li, Dongzhan Zhou

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, updated version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24063 2025-10-28 cs.CL cs.DB 79%

TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine

Jiacheng Xie, Yang Yu, Ziyang Zhang, Shuai Zeng, Jiaxuan He, Ayush Vasireddy, Xiaoting Tang, Congyu Guo, Lening Zhao, Congcong Jing, Guanghui An, Dong Xu

机构 * Department of Electrical Engineering and Computer Science, University of Missouri(密苏里大学电气工程与计算机科学系) Christopher S. Bond Life Sciences Center, University of Missouri(密苏里大学克里斯托弗·S·邦德生命科学中心) Department of Computer Science, Northwestern University(西北大学计算机科学系) Department of Computer Science and Mathematics, Truman State University(特拉华州立大学计算机科学与数学系) Marquette High School(马基特高中) Community Health Service Center, Shanghai Pudong New Area(上海浦东新区社区卫生服务中心) School of Engineering and Applied Science, University of Pennsylvania(宾夕法尼亚大学工程与应用科学学院) Department of Endocrinology, Seventh People’s Hospital of Shanghai University of Traditional Chinese Medicine(上海中医药大学第七人民医院内分泌科) School of Acupuncture-Moxibustion and Tuina, Shanghai University of Traditional Chinese Medicine(上海中医药大学针灸推拿学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22987 2025-10-28 cs.CE 78%

Capsule Network-Based Multimodal Fusion for Mortgage Risk Assessment from Unstructured Data Sources

Mahsa Tavakoli, Rohitash Chandra, Cristian Bravo

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23325 2025-10-28 cs.CV cs.AI cs.LG 76%

Multitask Multimodal Self-Supervised Learning for Medical Images

Cristian Simionescu

机构 * Department of Computer Science(计算机科学系) "Alexandru Ioan Cuza" University(阿莱克桑德鲁·伊奥安·库扎大学)

专题命中 多模态评测 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21881 2025-10-28 cs.AI cs.CL 73%

GeoThought: A Dataset for Enhancing Mathematical Geometry Reasoning in Vision-Language Models

Nannan Shi, Chuanyu Qin, Shipeng Song, Man Luo

机构 * Baidu Inc.(百度公司) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Intel Lab, Intel(英特尔实验室)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22571 2025-10-28 cs.CV cs.AI cs.MM 67%

STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models

Mahiro Ukai, Shuhei Kurita, Nakamasa Inoue

机构 * Institute of Science Tokyo(东京科学研究所) National Institute of Informatics(日本信息处理学会)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10281 2025-10-28 cs.CR cs.AI cs.CL cs.CV cs.LG 67%

ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test

Guan-Yan Yang, Tzu-Yu Cheng, Ya-Wen Teng, Farn Wanga, Kuo-Hui Yeh

机构 * Department of Electrical Engineering, National Taiwan University(国立台湾大学电子工程系) GARMIN (ASIA) CORPORATION(GARMIN(亚洲)公司) Institute of Artificial Intelligence Innovation, National Yang Ming Chiao Tung University(国家阳明交通大学人工智能创新研究所) Department of Information Management, National Dong Hwa University(国立东吴大学资讯管理系)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 30 pages, 22 figures. This preprint has been accepted for publication in Elsevier JOURNAL OF NETWORK AND COMPUTER APPLICATIONS (JNCA)

Journal ref Journal of Network and Computer Applications, Vol. 244, (2025) 104356

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23482 2025-10-28 cs.CV cs.AI 62%

On the Faithfulness of Visual Thinking: Measurement and Enhancement

Zujing Liu, Junwen Pan, Qi She, Yuan Gao, Guisong Xia

机构 * Wuhan University(武汉大学) ByteDance(字节跳动)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22760 2025-10-28 eess.IV cs.CV cs.MM 62%

Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions

Kai Ye, Bowen Liu, Jianghang Lin, Jiayi Ji, Pingyang Dai, Liujuan Cao

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China(教育部多媒体可信感知与高效计算重点实验室) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院) Xiamen University(厦门大学) School of Informatics, Xiamen University(厦门大学信息学院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01737 2025-10-28 cs.CV cs.MM 62%

Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark

Bingchen Miao, Wenqiao Zhang, Juncheng Li, Wangyu Wu, Siliang Tang, Zhaocheng Li, Haochen Shi, Jun Xiao, Yueting Zhuang

机构 * Zhejiang University(浙江大学) University of Liverpool(利物浦大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21810 2025-10-28 cs.CV cs.AI 62%

Hybrid Deep Learning Framework for Enhanced Diabetic Retinopathy Detection: Integrating Traditional Features with AI-driven Insights

Arpan Maity, Aviroop Pal, MD. Samiul Islam, Tamal Ghosh

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21807 2025-10-28 cs.CV cs.AI 62%

Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs

Jiaao Yu, Shenwei Li, Mingjie Han, Yifei Yin, Wenzheng Song, Chenghao Jia, Man Lan

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22237 2025-10-28 eess.AS cs.LG 57%

Bridging the Perceptual-Statistical Gap in Dysarthria Assessment: Why Machine Learning Still Falls Short

Krishna Gurugubelli

机构 * Samsung Research \& Development Institute Bengaluru, India\ .

专题命中 多模态评测 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22803 2025-10-28 cs.CV 57%

MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering

Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le

机构 * Faculty Of Information Technology VNU University of Engineering(信息科技学院越南工程大学) IT-BT Convergence Technology Division Vietnam-Korea Institute of Science(IT-BT融合技术部门越南-韩国科学技术院) TADI Global Lab TADI Global Company Limited(TADI全球实验室TADI全球有限公司) Faculty of Finance Banking Academy of Vietnam(金融学院越南银行学院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 4 figures, IEEE conference format

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22716 2025-10-28 cs.CV 57%

LRW-Persian: Lip-reading in the Wild Dataset for Persian Language

Zahra Taghizadeh, Mohammad Shahverdikondori, Arian Noori, Alireza Dadgarnia

机构 * Department of Mechanical Engineering(机械工程系) Sharif University of Technology(谢里夫科技大学) Department of Mathematical Sciences(数学科学系) Department of Computer Engineernig(计算机工程系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22225 2025-10-28 cs.CV 57%

Audio Frequency-Time Dual Domain Evaluation on Depression Diagnosis

Yu Luo, Nan Huang, Sophie Yu, Hendry Xu, Jerry Wang, Colin Wang, Zhichao Liu, Chen Zeng

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11147 2025-10-28 cs.CV 57%

3D-RAD: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks

Xiaotang Gai, Jiaxiang Liu, Yichen Li, Zijie Meng, Jian Wu, Zuozhu Liu

机构 * ZJU-Angelalign R&D Center for Intelligence Healthcare(浙江大学智能医疗研发中心) Zhejiang University(浙江大学) Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence(浙江省医学影像人工智能重点实验室) Guangdong Institute of Intelligence Science and Technology(广东省智能科学与技术研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21801 2025-10-28 cs.CV cs.LG 57%

Morphology-Aware KOA Classification: Integrating Graph Priors with Vision Models

Marouane Tliba, Mohamed Amine Kerkouri, Yassine Nasser, Nour Aburaed, Aladine Chetouani, Ulas Bagci, Rachid Jennane

机构 * University of Orleans(奥尔良大学) University of Sorbonne Paris Nord(巴黎-萨特大学) Northwestern University(西北大学) University of Dubai(迪拜大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22274 2025-10-28 cs.CR cs.LG 50%

SecureLearn -- An Attack-agnostic Defense for Multiclass Machine Learning Against Data Poisoning Attacks

Anum Paracha, Junaid Arshad, Mohamed Ben Farah, Khalid Ismail

机构 * College of Computing, Birmingham City University, Birmingham, United Kingdom(计算学院,伯明翰城市大学,伯明翰,英国)

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18936 2025-10-28 cs.IR cs.SE 50%

SBAN: A Framework & Multi-Dimensional Dataset for Large Language Model Pre-Training and Software Code Mining

Hamed Jelodar, Mohammad Meymani, Samita Bai, Roozbeh Razavi-Far, Ali A. Ghorbani

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22800 2025-10-28 cs.NE 50%

Probing the Representational Geometry of Color Qualia: Dissociating Pure Perception from Task Demands in Brains and AI Models

Jing Xu

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22498 2025-10-28 cs.HC 50%

Emotion Recognition with Minimal Wearable Sensing: Multi-domain Feature, Hybrid Feature Selection, and Personalized vs. Generalized Ensemble Model Analysis

Muhammad Irfan, Anum Nawaz, Ayse Kosal Bulbul, Riku Klen, Abdulhamit Subasi, Tomi Westerlund, Wei Chen

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏