arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-11 至 2025-09-11 共收录 38 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2509.08742 2025-09-11 q-fin.CP cs.AI 83%

FinZero: Launching Multi-modal Financial Time Series Forecast with Large Reasoning Model

Yanlong Wang, Jian Xu, Fei Ma, Hongkang Zhang, Hang Yu, Tiantian Gao, Yu Wang, Haochen You, Shao-Lun Huang, Danny Dongning Sun, Xiao-Ping Zhang

机构 * Tsinghua University(清华大学) Pengcheng Laboratory(鹏城实验室) Guangming Laboratory(光明实验室) Ant Group(蚂蚁集团) Columbia University(哥伦比亚大学) Southern University of Science and Technology(南方科技大学)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08715 2025-09-11 cs.CV 83%

BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion

Sike Xiang, Shuang Chen, Amir Atapour-Abarghouei

机构 * Department of Computer Science, Durham University(计算机科学系,杜伦大学)

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10160 2025-09-11 cs.CV cs.AI 81%

PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval

Jiancheng Pan, Muyuan Ma, Qing Ma, Cong Bai, Shengyong Chen

机构 * IEEE Publication Technology Department(IEEE出版技术部)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08618 2025-09-11 cs.CV 79%

CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging

Zhihao Zhao, Yinzheng Zhao, Junjie Yang, Xiangtong Yao, Quanmin Liang, Shahrooz Faghihroohi, Kai Huang, Nassir Navab, M. Ali Nasseri

机构 * Technical University of Munich(慕尼黑技术大学) Sun Yat-Sen University(中山大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments BIBM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08570 2025-09-11 cs.CV 70%

Vision-Language Semantic Aggregation Leveraging Foundation Model for Generalizable Medical Image Segmentation

Wenjun Yu, Yinchen Zhou, Jia-Xuan Jiang, Shubin Zeng, Yuee Li, Zhong Wang

机构 * organization= School of Information Science \& Engineering, Lanzhou University , addressline= , city= Lanzhou , postcode= 730000 , country= China

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 29 pages and 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08490 2025-09-11 cs.CV cs.AI 62%

A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models

Edwine Nabahirwa, Wei Song, Minghua Zhang, Yi Fang, Zhou Ni

机构 * Shanghai Ocean University(上海海洋大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 72 Pages, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06130 2025-09-11 cs.CV cs.CL 62%

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

Ce Zhang, Zifu Wan, Zhehan Kan, Martin Q. Ma, Simon Stepputtis, Deva Ramanan, Russ Salakhutdinov, Louis-Philippe Morency, Katia Sycara, Yaqi Xie

机构 * School of Computer Science, Carnegie Mellon University(计算机科学系,卡内基梅隆大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICLR 2025. Project page: https://zhangce01.github.io/DeGF/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15108 2025-09-11 cs.LG cs.AI cs.RO 57%

VIPER: Visual Perception and Explainable Reasoning for Sequential Decision-Making

Mohamed Salim Aissi, Clemence Grislain, Mohamed Chetouani, Olivier Sigaud, Laure Soulier, Nicolas Thome

机构 * Sorbonne Université, CNRS, ISIR(索邦大学、国家科学研究中心、ISIR)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03521 2025-09-11 cs.CV 57%

Have Large Vision-Language Models Mastered Art History?

Ombretta Strafforello, Derya Soydaner, Michiel Willems, Anne-Sofie Maerten, Stefanie De Winter

机构 * KU Leuven(库勒韦恩大学) Leiden University(莱顿大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 2 篇

2509.08689 2025-09-11 cs.HC 78%

Augmenting speech transcripts of VR recordings with gaze, pointing, and visual context for multimodal coreference resolution

Riccardo Bovo, Frederik Brudy, George Fitzmaurice, Fraser Anderson

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07538 2025-09-11 cs.CV 57%

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Peijin Xie, Shun Qian, Bingquan Liu, Dexin Wang, Lin Sun, Xiangzheng Zhang

机构 * IEEE

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 2 篇

2508.14581 2025-09-11 cs.MM eess.IV 79%

Memory-Anchored Multimodal Reasoning for Explainable Video Forensics

Chen Chen, Runze Li, Zejun Zhang, Pukun Zhao, Fanqing Zhou, Longxiang Wang, Haojian Huang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08354 2025-09-11 cs.RO cs.AI 57%

Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration

Ce Guo, Xieyuanli Chen, Zhiwen Zeng, Zirui Guo, Yihong Li, Haoran Xiao, Dewen Hu, Huimin Lu

机构 * College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments 20 pages, 19 figures, accepted by IEEE Transactions on Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2509.08216 2025-09-11 cs.IR 82%

Vector embedding of multi-modal texts: a tool for discovery?

Beth Plale, Sai Navya Jyesta, Sachith Withana

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08338 2025-09-11 cs.CV cs.AI cs.LG 81%

Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis

Jihyun Moon, Charmgil Hong

机构 * Handong Global University(-handong全球大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Medical Image Computing and Computer-Assisted Intervention (MICCAI) ISIC Skin Image Analysis Workshop (MICCAI ISIC) 2025; 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08624 2025-09-11 cs.CV cs.AI 62%

UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation

Zhihao Zhao, Yinzheng Zhao, Junjie Yang, Xiangtong Yao, Quanmin Liang, Daniel Zapp, Kai Huang, Nassir Navab, M. Ali Nasseri

机构 * Technical University of Munich(慕尼黑技术大学) Sun Yat-Sen University(中山大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments BIBM

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 5 篇

2509.08519 2025-09-11 cs.CV cs.MM 84%

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Liyang Chen, Tianxiang Ma, Jiawei Liu, Bingchuan Li, Zhuowei Chen, Lijie Liu, Xu He, Gen Li, Qian He, Zhiyong Wu

机构 * Tsinghua University(清华大学) Intelligent Creation Lab, ByteDance(字节跳动智能创作实验室)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08489 2025-09-11 cs.CV cs.AI 81%

Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation

Kaleem Ahmad

机构 * Independent Researcher(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08442 2025-09-11 cs.CV cs.AI cs.LG q-bio.NC 62%

Spherical Brownian Bridge Diffusion Models for Conditional Cortical Thickness Forecasting

Ivan Stoyanov, Fabian Bongratz, Christian Wachinger

机构 * Lab for AI in Medical Imaging, Technical University of Munich, Munich, Germany(人工智能医学影像实验室,慕尼黑技术大学,慕尼黑,德国) Munich Center for Machine Learning, Munich, Germany(慕尼黑机器学习中心,慕尼黑,德国)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06932 2025-09-11 cs.RO cs.CV 57%

LLaDA-VLA: Vision Language Diffusion Action Models

Yuqing Wen, Hebei Li, Kefan Gu, Yucheng Zhao, Tiancai Wang, Xiaoyan Sun

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学) Dexmal Project Page(Dexmal项目页)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04242 2025-09-11 cs.LG cs.CV 57%

Task-based Loss Functions in Computer Vision: A Comprehensive Review

Omar Elharrouss, Yasir Mahmood, Yassine Bechqito, Mohamed Adel Serhani, Elarbi Badidi, Jamal Riffi, Hamid Tairi

机构 * Department of Computer Science and Software Engineering, College of Information Technology, United Arab Emirates University.(计算机科学与软件工程系,信息科技学院,阿联酋大学) Department of Information Systems, College of Computing and Informatics, University of Sharjah, Sharjah, United Arab Emirates(信息系统系,计算与信息学院,沙迦大学) Department of Informatics, Faculty of Sciences Dhar El Mahraz, Sidi Mohamed Ben Abdellah University, Fez, Morocco(信息学系,达尔·埃尔·马哈勒兹学院,西迪·莫哈梅德·本·阿卜杜勒拉赫曼大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 10 篇

2509.08777 2025-09-11 cs.CV cs.CL 87%

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

Eric Slyman, Mehrab Tanjim, Kushal Kafle, Stefan Lee

机构 * Adobe(Adobe公司) Oregon State University(俄勒冈州立大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(title);分类 cs.CV、cs.CL

Comments 17 pages, 8 figures, Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08800 2025-09-11 cs.SD cs.AI cs.CV cs.MM eess.AS 85%

PianoVAM: A Multimodal Piano Performance Dataset

Yonghyun Kim, Junhyung Park, Joonhyung Bae, Kirak Kim, Taegyun Kwon, Alexander Lerch, Juhan Nam

专题命中 多模态评测 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to the 26th International Society for Music Information Retrieval (ISMIR) Conference, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08008 2025-09-11 cs.SI cs.AI cs.MM 84%

A New Dataset and Benchmark for Grounding Multimodal Misinformation

Bingjian Yang, Danni Xu, Kaipeng Niu, Wenxuan Liu, Zheng Wang, Mohan Kankanhalli

机构 * National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(多媒体软件国家工程研究中心,计算机科学学院,武汉大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学) School of Computer Science, Peking University(计算机科学学院,北京大学) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments 6 pages, 5 figures, ACM Multimedia 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01306 2025-09-11 cs.AI cs.CV 81%

Multimodal Medical Disease Classification with LLaMA II

Christian Gapp, Elias Tappeiner, Martin Welk, Rainer Schubert

机构 * Institute of Biomedical Image Analysis UMIT TIROL(生物医学影像分析研究所 UMIT TIROL) UMIT TIROL – Private University for Health Sciences and Health Technology(UMIT TIROL – 健康科学与健康技术私立大学) VASCage – Centre on Clinical Stroke Research Innsbruck(VASCage – 神经科学临床研究中心 Innsbruck)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, conference: AIRoV -- The First Austrian Symposium on AI, Robotics, and Vision 25.-27.3.2024, Innsbruck

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08694 2025-09-11 cs.CV 79%

Multi-Modal Robust Enhancement for Coastal Water Segmentation: A Systematic HSV-Guided Framework

Zhen Tian, Christos Anagnostopoulos, Qiyuan Wang, Zhiwei Gao

机构 * School of Computing Science, University of Glasgow(格拉斯哥大学计算科学学院) School of Computing Engineering, University of Glasgow(格拉斯哥大学计算工程学院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08024 2025-09-11 cs.CV cs.CY 79%

Two Stage Context Learning with Large Language Models for Multimodal Stance Detection on Climate Change

Lata Pangtey, Omkar Kabde, Shahid Shafi Dar, Nagendra Kumar

机构 * Department of Computer Science and Engineering, Indian Institute of Technology (IIT) Indore(计算机科学与工程系,印度理工学院(IIT)印多尔) Chaitanya Bharathi Institute of Technology, Gandipet, Hyderabad, 500075(恰伊坦尼亚·巴哈提技术学院,加迪皮特,海得拉巴,500075)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07994 2025-09-11 eess.IV cs.CV cs.LG 74%

STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery

David Robinson, Animesh Gupta, Rizwan Quershi, Qiushi Fu, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,佛罗里达中央大学) Mechanical and Aerospace Engineering, University of Central Florida(机械与航空航天工程,佛罗里达中央大学)

专题命中 多模态评测 :multimodal(title);分类 cs.CV

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16822 2025-09-11 cs.CV cs.AI 73%

Integrating Clinical Knowledge Graphs and Gradient-Based Neural Systems for Enhanced Melanoma Diagnosis via the 7-Point Checklist

Yuheng Wang, Tianze Yu, Jiayue Cai, Sunil Kalia, Harvey Lui, Z. Jane Wang, Tim K. Lee

机构 * The University of British Columbia(不列颠哥伦比亚大学) Vancouver Coastal Health Research Institute(温哥华海岸健康研究机构) Department of Dermatology and Skin Science(皮肤科与皮肤病学系) Photomedicine Institute(光医学研究所) Centre for Clinical Epidemiology and Evaluation(临床流行病学与评估中心) School of Biomedical Engineering(生物医学工程学院) Department of Electrical and Computer Engineering(电气与计算机工程系) Department of Population Health Sciences(人群健康科学系) BC Cancer(不列颠哥伦比亚省癌症中心) BC Children’s Hospital Research Institute(不列颠哥伦比亚儿童医院研究机构) Guangdong Key Laboratory of Biomedical Measurements and Ultrasound Imaging(广东生物医学测量与超声成像重点实验室) Shenzhen University Medical School(深圳大学医学院) Shenzhen University(深圳大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments The paper was officially accepted for publication in IEEE Transactions on Neural Networks and Learning Systems in August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13729 2025-09-11 cs.CL 57%

Baba Is AI: Break the Rules to Beat the Benchmark

Nathan Cloos, Meagan Jens, Michelangelo Naim, Yen-Ling Kuo, Ignacio Cases, Andrei Barbu, Christopher J. Cueva

机构 * MIT(麻省理工学院) Department of Computer Science, University of Virginia, USA(弗吉尼亚大学计算机科学系)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏