arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-18 至 2025-11-18 共收录 142 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 30 篇

2511.11597 2025-11-18 cs.AI cs.CL 62%

CLINB: A Climate Intelligence Benchmark for Foundational Models

Michelle Chen Huebscher, Katharine Mach, Aleksandar Stanić, Markus Leippold, Ben Gaiarin, Zeke Hausfather, Elisa Rawat, Erich Fischer, Massimiliano Ciaramita, Joeri Rogelj, Christian Buck, Lierni Sestorain Saralegui, Reto Knutti

机构 * University of Miami(迈阿密大学) University of Zurich(苏黎世大学) Stripe(Stripe公司) ETH Zurich(苏黎世联邦理工学院) Imperial College London(伦敦帝国理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Questions, system prompt and model judge prompts available here: https://www.kaggle.com/datasets/deepmind/clinb-questions

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13626 2025-11-18 cs.AI 57%

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

Kaiwen Xue, Chenglong Li, Zhonghong Ou, Guoxin Zhang, Kaoyan Lu, Shuai Lyu, Yifan Zhu, Ping Zong Junpeng Ding, Xinyu Liu, Qunlin Chen, Weiwei Qin, Yiran Shen, Jiayi Cen

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 13 pages, 3 figures,The 40th Annual AAAI Conference on Artificial Intelligence(AAAI 2026),Paper has been accepted for a poster presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13381 2025-11-18 cs.CL 57%

Can Large Language Models Function as Qualified Pediatricians? A Systematic Evaluation in Real-World Clinical Contexts

Siyu Zhu, Mouxiao Bian, Yue Xie, Yongyu Tang, Zhikang Yu, Tianbin Li, Pengcheng Chen, Bing Han, Jie Xu, Xiaoyan Dong

机构 * Shanghai Children’s Hospital, School of Medicine,Shanghai Jiao Tong University(上海儿童医学中心,上海交通大学医学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai University of Traditional Chinese Medicine(上海中医药大学) University of Washington(华盛顿大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12928 2025-11-18 cs.CL 57%

Visual Room 2.0: Seeing is Not Understanding for MLLMs

Haokun Li, Yazhou Zhang, Jizhi Ding, Qiuchi Li, Peng Zhang

机构 * Tianjin University(天津大学) Shandong Institute of Petroleum and Chemical Technology(山东石油化学技术学院) Beijing Institute of Technology(北京理工大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12648 2025-11-18 cs.CR cs.AI cs.LG 57%

Scalable Hierarchical AI-Blockchain Framework for Real-Time Anomaly Detection in Large-Scale Autonomous Vehicle Networks

Rathin Chandra Shit, Sharmila Subudhi

机构 * organization= Dept. of Computer Science \& Engg., International Institute of Information Technology , city= Bhubaneswar , postcode= 751003 , state= Odisha , country= India organization= Dept. of Computer Science, Maharaja Sriram Chandra Bhanja Deo University , city= Baripada , postcode= 757003 , state= Odisha , country= India

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments Submitted to the Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02206 2025-11-18 cs.CV 57%

Language-Enhanced Generative Modeling for Amyloid PET Synthesis from MRI and Blood Biomarkers

Zhengjie Zhang, Xiaoxie Mao, Qihao Guo, Shaoting Zhang, Qi Huang, Mu Zhou, Fang Xie, Mianxin Liu

机构 * organization= Shanghai Artificial Intelligence Laboratory , city= Shanghai , postcode= 200082 , country= China organization= School of Medicine, Xiamen University , city= Xiamen , state= Fujian , country= China organization= Department of Nuclear Medicine \& PET Center, Huashan Hospital, Fudan University , city= Shanghai , country= China organization= Department of Gerontology, Shanghai Jiao Tong University Affiliated Sixth People’s Hospital , city= Shanghai , country= China organization= Department of Computer Science, Rutgers University , city= New Brunswick , state= New Jersey , country= United States organization= Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences , city= Shenzhen , country= China

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 31 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12026 2025-11-18 cs.CV 57%

Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and Benchmark

Rulin Zhou, Wenlong He, An Wang, Jianhang Zhang, Xuanhui Zeng, Xi Zhang, Chaowei Zhu, Haijun Hu, Hongliang Ren

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments AAAI 2026 oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13343 2025-11-18 cs.AR 50%

Coliseum project: Correlating climate change data with the behavior of heritage materials

A Cormier, David Roqui, Fabrice Surma, Martin Labouré, Jean-Marc Vallet, Odile Guillon, N Grozavu, Ann Bourgès

专题命中 多模态评测 :multimodal(abstract)

Journal ref Stone 2025 : 15th International Congress on the Deterioration and Conservation of Stone, Sep 2025, Paris, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12990 2025-11-18 q-bio.QM 50%

Brain Networks Flow-Topology via Variance Minimization in the Wasserstein Space

Sixtus Dakurah

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10213 2025-11-18 cs.LG 50%

Out-of-Context Misinformation Detection via Variational Domain-Invariant Learning with Test-Time Training

Xi Yang, Han Zhang, Zhijian Lin, Yibiao Hu, Hong Han

专题命中 多模态评测 :image-text(abstract)

Comments accepted by the AAAI Conference on Artificial Intelligence (AAAI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00730 2025-11-18 cs.HC 50%

Teaching LLMs to See and Guide: Context-Aware Real-Time Assistance in Augmented Reality

Mahya Qorbani, Kamran Paynabar, Mohsen Moghaddam

专题命中 多模态评测 :multimodal(abstract)

Comments This work has been submitted to the IEEE Transactions on Systems, Man, and Cybernetics: Systems for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07083 2025-11-18 cs.NE 50%

A Generalized and Configurable Benchmark Generator for Continuous Unconstrained Numerical Optimization

Amir H. Gandomi, Mohammad Nabi Omidvar, Rohit Salgotra, Kalyanmoy Deb

专题命中 多模态评测 :multimodal(abstract)

Comments 14 pages, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 7 篇

2511.13476 2025-11-18 cs.AI 79%

Multi-Agent Multimodal Large Language Model Framework for Automated Interpretation of Fuel Efficiency Analytics in Public Transportation

Zhipeng Ma, Ali Rida Bahja, Andreas Burgdorf, André Pomp, Tobias Meisen, Bo Nørregaard Jørgensen, Zheng Grace Ma

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Journal ref Applied Sciences, 2025, 15(21), 11619

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12586 2025-11-18 cs.CL 79%

MMWOZ: Building Multimodal Agent for Task-oriented Dialogue

Pu-Hai Yang, Heyan Huang, Heng-Da Xu, Fanshu Sun, Xian-Ling Mao, Chaoxu Mu

机构 * School of Artificial Intelligence, Anhui University, Hefei, China(人工智能学院,安徽大学,合肥,中国) School of Computer Science(计算机科学学院) School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China(计算机科学与技术学院,北京理工大学,北京,中国)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19768 2025-11-18 cs.CL 79%

T^2Agent A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search

Xing Cui, Yueying Zou, Zekun Li, Peipei Li, Xinyuan Xu, Xuannan Liu, Huaibo Huang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

Comments accepted by AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11751 2025-11-18 cs.CV cs.AI cs.MA 62%

Concept-RuleNet: Grounded Multi-Agent Neurosymbolic Reasoning in Vision Language Models

Sanchit Sinha, Guangzhi Xiong, Zhenghao He, Aidong Zhang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments AAAI 2026 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04614 2025-11-18 cs.AI 57%

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

Yuyang Wanyan, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Jiabo Ye, Yutong Kou, Ming Yan, Fei Huang, Xiaoshan Yang, Weiming Dong, Changsheng Xu

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences, China(自动化研究所,中国科学院,中国) School of Artificial Intelligence, University of Chinese Academy of Sciences, China(中国科学院大学人工智能学院,中国) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11120 2025-11-18 cs.CR 50%

SoK: How Sensor Attacks Disrupt Autonomous Vehicles: An End-to-end Analysis, Challenges, and Missed Threats

Qingzhao Zhang, Shaocheng Luo, Z. Morley Mao, Miroslav Pajic, Michael K. Reiter

专题命中 多模态Agent :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19838 2025-11-18 cs.HC 50%

LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects

Guangyi Liu, Pengxiang Zhao, Yaozhen Liang, Liang Liu, Yaxuan Guo, Han Xiao, Weifeng Lin, Yuxiang Chai, Yue Han, Shuai Ren, Hao Wang, Xiaoyu Liang, WenHao Wang, Tianze Wu, Zhengxi Lu, Siheng Chen, LiLinghao, Hao Wang, Guanjing Xiong, Yong Liu, Hongsheng Li

专题命中 多模态Agent :multimodal(abstract)

Comments Paper accepted to TMLR 2025, Project Homepage: https://github.com/PhoneLLM/Awesome-LLM-Powered-Phone-GUI-Agents

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 25 篇

2511.12982 2025-11-18 cs.CR cs.CV 83%

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

Xuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou, Bo Du, Mang Ye

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17686 2025-11-18 cs.CV 83%

Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration

Yuhang Han, Xuyang Liu, Zihan Zhang, Pengxiang Ding, Junjie Chen, Donglin Wang, Honggang Chen, Qingsen Yan, Siteng Huang

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00374 2025-11-18 cs.CV cs.AI cs.MM 82%

MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention

Tianyi Wang, Jianan Fan, Dingxin Zhang, Dongnan Liu, Yong Xia, Heng Huang, Weidong Cai

机构 * The University of Sydney(悉尼大学) School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) School of Computer Science and Engineering, Northwestern Polytechnical University(西北工业大学计算机科学与工程学院) University of Maryland(马里兰大学) Ningbo Institute of Northwestern Polytechnical University(西北工业大学宁波学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by IEEE Transactions on Medical Imaging (TMI). Code available at https://github.com/TianyiFranklinWang/MIRROR. Project page: https://tianyifranklinwang.github.io/MIRROR

Journal ref IEEE Trans. Med. Imaging (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10218 2025-11-18 cs.AI 79%

MTP: Exploring Multimodal Urban Traffic Profiling with Modality Augmentation and Spectrum Fusion

Haolong Xiang, Peisi Wang, Xiaolong Xu, Kun Yi, Xuyun Zhang, Quanzheng Sheng, Amin Beheshti, Wei Fan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13637 2025-11-18 cs.LG 79%

Towards Multimodal Representation Learning in Paediatric Kidney Disease

Ana Durica, John Booth, Ivana Drobnjak

机构 * Institute of Health Informatics(健康信息学研究所) University College London(伦敦大学学院) Data Research, Innovation and Virtual Environments Unit(数据研究、创新与虚拟环境单位) Great Ormond Street Hospital(格雷特奥蒙德医院) Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 4 pages, 3 figures. EurIPS 2025 Multimodal Representation Learning for Healthcare (MMRL4H) workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10962 2025-11-18 cs.IR 78%

LEMUR: Large scale End-to-end MUltimodal Recommendation

Xintian Han, Honggang Chen, Quan Lin, Jingyue Gao, Xiangyuan Ren, Lifei Zhu, Zhisheng Ye, Shikang Wu, XiongHang Xie, Xiaochu Gan, Bingzheng Wei, Peng Xu, Zhe Wang, Yuchao Zheng, Jingjian Lin, Di Wu, Junfeng Ge

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10644 2025-11-18 cs.LG 78%

Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection

Yihua Wang, Qi Jia, Cong Xu, Feiyu Chen, Yuhan Liu, Haotian Zhang, Liang Jin, Lu Liu, Zhichun Wang

机构 * IEIT SYSTEMS Co., Ltd.(IEIT SYSTEMS公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at AAAI 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12917 2025-11-18 cs.CV 77%

Explore How to Inject Beneficial Noise in MLLMs

Ruishu Zhu, Sida Huang, Ziheng Jiao, Hongyuan Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13283 2025-11-18 cs.CV 70%

TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing

Jongha Kim, Minseong Bae, Sanghyeok Lee, Jinsung Yoon, Hyunwoo J. Kim

机构 * Korea University(韩国大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments AAAI 2026 (Main Technical Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13575 2025-11-18 cs.CV cs.AI 62%

Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification

Linhan Zhou, Shuang Li, Neng Dong, Yonghang Tai, Yafei Zhang, Huafeng Li

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures, accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11730 2025-11-18 cs.CV cs.AI 62%

GROVER: Graph-guided Representation of Omics and Vision with Expert Regulation for Adaptive Spatial Multi-omics Fusion

Yongjun Xiao, Dian Meng, Xinlei Huang, Yanran Liu, Shiwei Ruan, Ziyue Qiao, Xubin Zheng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 3 figures, Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏