arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-23 至 2025-10-23 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2506.16962 2025-10-23 cs.CV cs.AI cs.CL 85%

Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search

Haoran Sun, Yankai Jiang, Wenjie Lou, Yujie Zhang, Wenjie Li, Lilong Wang, Mianxin Liu, Lei Liu, Xiaosong Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19451 2025-10-23 cs.CV cs.MM 81%

Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis

Xueqi Ma, Yanbei Jiang, Sarah Erfani, James Bailey, Weifeng Liu, Krista A. Ehinger, Jey Han Lau

机构 * The University of Melbourne(墨尔本大学) China University of Petroleum (East China)(中国石油大学(华东))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19183 2025-10-23 cs.CV cs.AI 81%

PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning

Fengyuan Sun, Hui Chen, Xinhao Xu, Dandan Zheng, Jingdong Chen, Jun Zhou, Jungong Han, Guiguang Ding

机构 * School of Software, Tsinghua University(清华大学软件学院) Ant Group(蚂蚁集团) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00711 2025-10-23 cs.LG cs.AI cs.CV 81%

QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training

Wei Dai, Peilin Chen, Chanakya Ekbote, Paul Pu Liang

机构 * MIT Media Lab(MIT媒体实验室) MIT EECS(MIT电子工程与计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted as Oral at NeurIPS 2025. Revision after camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19305 2025-10-23 cs.LG cs.CV 79%

FrogDeepSDM: Improving Frog Counting and Occurrence Prediction Using Multimodal Data and Pseudo-Absence Imputation

Chirag Padubidri, Pranesh Velmurugan, Andreas Lanitis, Andreas Kamilaris

机构 * Pervasive Systems, University of Twente(普罗威斯系统,埃因霍温大学) CYENS Center of Excellence(CYENS卓越中心) Cyprus University of Tecnhology(塞浦路斯技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18582 2025-10-23 cs.CV 79%

The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers

Daiqing Qi, Handong Zhao, Jing Shi, Simon Jenni, Yifei Fan, Franck Dernoncourt, Scott Cohen, Sheng Li

机构 * University of Virginia(弗吉尼亚大学) Adobe(Adobe公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19339 2025-10-23 cs.CV cs.CL 73%

PixelWorld: How Far Are We from Perceiving Everything as Pixels?

Zhiheng Lyu, Xueguang Ma, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) Vector Institute, Toronto(多伦多向量研究所)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19599 2025-10-23 cs.CV cs.AI 62%

XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography

Haozhe Luo, Shelley Zixin Shu, Ziyu Zhou, Sebastian Otalora, Mauricio Reyes

机构 * ARTORG Center for Biomedical Engineering Research, University of Bern, Switzerland(ARTORG生物医学工程研究中心,伯尔尼大学,瑞士) Shanghai Jiao Tong University, China(上海交通大学,中国) Kaiko.AI, Switzerland(Kaiko.AI,瑞士) Dept. of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,因斯普尔茨医院,伯尔尼大学医院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05571 2025-10-23 cs.CL 57%

Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations

Chengzhi Liu, Yuzhe Yang, Kaiwen Zhou, Zhen Zhang, Yue Fan, Yanan Xie, Peng Qi, Xin Eric Wang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09024 2025-10-23 q-bio.NC cs.CV 57%

CNeuroMod-THINGS, a densely-sampled fMRI dataset for visual neuroscience

Marie St-Laurent, Basile Pinsard, Oliver Contier, Elizabeth DuPre, Katja Seeliger, Valentina Borghesani, Julie A. Boyle, Lune Bellec, Martin N. Hebart

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 16 pages manuscript, 5 figures, 9 pages supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10043 2025-10-23 cs.SE 50%

TrioXpert: An Automated Incident Management Framework for Microservice System

Yongqian Sun, Yu Luo, Xidao Wen, Yuan Yuan, Xiaohui Nie, Shenglin Zhang, Tong Liu, Xi Luo

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏