arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2509.20022 2025-09-25 cs.CV 83%

PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival Prediction

Manahil Raza, Ayesha Azam, Talha Qaiser, Nasir Rajpoot

机构 * University of Warwick, UK(沃里克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19628 2025-09-25 cs.CE cs.CL q-fin.CP 83%

Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series

Ross Koval, Nicholas Andrews, Xifeng Yan

机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12623 2025-09-24 cs.SD cs.AI cs.CL cs.MM eess.AS 83%

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji

机构 * Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted to EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18221 2025-09-24 cs.AI cs.LG 83%

Multimodal Health Risk Prediction System for Chronic Diseases via Vision-Language Fusion and Large Language Models

Dingxin Lu, Shurui Wu, Xinyi Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16149 2025-09-22 cs.CV 83%

Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models

Renjie Pi, Kehao Miao, Li Peihang, Runtao Liu, Jiahui Gao, Jipeng Zhang, Xiaofang Zhou

机构 * HKUST(香港科技大学) HKU(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16017 2025-09-22 cs.CV 83%

DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching

Meng Yang, Fan Fan, Zizhuo Li, Songchu Deng, Yong Ma, Jiayi Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14383 2025-09-19 cs.RO cs.CV 83%

RLBind: Adversarial-Invariant Cross-Modal Alignment for Unified Robust Embeddings

Yuhong Lu

机构 * Samueli School of Engineering, Electrical and Computer Engineering, UCLA(UCLA电气与计算机工程学院萨姆利学校)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments This paper is submitted to IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06641 2025-09-16 cs.AI cs.LG 83%

CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning

Zhou-Peng Shou, Zhi-Qiang You, Fang Wang, Hai-Bo Liu

机构 * NoDesk AI Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :omni-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09427 2025-09-12 cs.CV 83%

FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

Yuchan Jie, Yushen Xu, Xiaosong Li, Fuqiang Zhou, Jianming Lv, Huafeng Li

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Physics and Optoelectronic Engineering, Foshan University(佛山大学物理与光电工程学院) School of Instrumentation Science and Optelectronics Engineering, Beihang University(北航仪器科学与光电工程学院) School of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref Information Fusion, 2025, 121: 103146

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20454 2025-09-10 cs.LG cs.CL 83%

CoMMIT: Coordinated Multimodal Instruction Tuning

Xintong Li, Junda Wu, Tong Yu, Yu Wang, Xiang Chen, Jiuxiang Gu, Lina Yao, Julian McAuley, Jingbo Shang

机构 * University of California, San Diego(加州大学圣地亚哥分校) Adobe Research(Adobe研究) The University of New South Wales(新南威尔士大学) CSIRO’s Data61(澳大利亚联邦科学与工业研究组织的数据61)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06976 2025-09-10 cs.LG cs.AI 83%

A Knowledge-Guided Cross-Modal Feature Fusion Model for Local Traffic Demand Prediction

Lingyu Zhang, Pengfei Xu, Guobin Wu, Jian Liang, Ruiyang Dong, Yunhai Wang, Xuan Song

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06291 2025-09-09 cs.CV 83%

Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding

Jiangnan Xie, Xiaolong Zheng, Liang Zheng

机构 * College of Electronics and Information, Hangzhou Dianzi University(电子信息学院,杭州电子大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04938 2025-09-08 cs.MM 83%

An Emotion Recognition Framework via Cross-modal Alignment of EEG and Eye Movement Data

Jianlu Wang, Yanan Wang, Tong Liu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20309 2025-09-08 cs.CV 83%

Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs

Zitian Wang, Yue Liao, Kang Rong, Fengyun Rao, Yibo Yang, Si Liu

机构 * Beihang University(北航大学) National University of Singapore(国立新加坡大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09556 2025-09-05 cs.CL 83%

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou, Efthymios Georgiou, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos

机构 * National Technical University of Athens(希腊国家技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03433 2025-09-04 cs.CV 83%

Decoding Visual Neural Representations by Multimodal with Dynamic Balancing

Kaili sun, Xingyu Miao, Bing Zhai, Haoran Duan, Yang Long

机构 * Department of Computer Science, Durham University, UK.(计算机科学系,杜ham大学,英国) Computer and Information Sciences, Northumbria University, UK.(计算机与信息科学,北umbria大学,英国) Department of Automation, Tsinghua University, China.(自动化系,清华大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00622 2025-09-03 cs.AI cs.IR 83%

BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting

Shiqiao Zhou, Holger Schöner, Huanbo Lyu, Edouard Fouché, Shuo Wang

机构 * University of Birmingham(伯明翰大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06989 2025-08-29 cs.CR cs.CV 83%

Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application

Wenzhuo Xu, Zhipeng Wei, Xiongtao Sun, Zonghao Ying, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang, Quanchen Zou

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05849 2025-08-26 cs.CV 83%

ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt

Fanhu Zeng, Fei Zhu, Haiyang Guo, Xu-Yao Zhang, Cheng-Lin Liu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室) School of Artificial Intelligence, UCAS(人工智能学院) Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13460 2025-08-20 cs.CV 83%

Revisiting MLLM Token Technology through the Lens of Classical Visual Coding

Jinming Liu, Junyan Lin, Yuntao Wei, Kele Shao, Keda Tao, Jianguo Huang, Xudong Yang, Zhibo Chen, Huan Wang, Xin Jin

机构 * Eastern Institute of Technology, Ningbo, China(东部技术研究所) Westlake University(西湖大学) USTC(中国科学技术大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13072 2025-08-19 cs.AI 83%

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis

Yuting Zhang, Tiantian Geng, Luoying Hao, Xinxing Cheng, Alexander Thorley, Xiaoxia Wang, Wenqi Lu, Sandeep S Hothi, Lei Wei, Zhaowen Qiu, Dipak Kotecha, Jinming Duan

机构 * School of Computer Science, University of Birmingham, Birmingham, UK Department of Cardiovascular Sciences, University of Birmingham, Birmingham, UK NIHR Birmingham Biomedical Research Centre West Midlands NHS Secure Data Environment, University Hospitals Birmingham NHS Foundation Trust, Birmingham, UK Department of Computing Mathematics, Manchester Metropolitan University, Manchester, UK Department of Cardiology, Heart Lung Centre, Royal Wolverhampton NHS Trust, Wolverhampton, UK Department of Cardiovascular Surgery, The First Affiliated Hospital with Nanjing Medical University , Nanjing,China College of Computer Control Engineering, Northeast Forestry University, Harbin, China Julius Center, University Medical Center Utrecht, the Netherlands Data Sciences, University of Manchester, Manchester, UK

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12917 2025-08-19 cs.CV 83%

CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection with IoU Joint Prediction

Zhiwei Ning, Zhaojiang Liu, Xuanang Gao, Yifan Zuo, Jie Yang, Yuming Fang, Wei Liu

机构 * School of Automation and Intelligent Sensing & Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(自动化与智能感知学院及图像处理与模式识别研究所,上海交通大学) School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics(计算机与人工智能学院,江西财经大学) School of Automation and Intelligent Sensing & Institute of Image Processing and Pattern Recognition & Institute of Medical Robotics, Shanghai Jiao Tong University(自动化与智能感知学院及图像处理与模式识别研究所及医学机器人研究所,上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments The Paper is Accepted by TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07264 2025-08-18 cs.SI cs.AI 83%

FLUID: Flow-Latent Unified Integration via Token Distillation for Expert Specialization in Multimodal Learning

Van Duc Cuong, Ta Dinh Tam, Tran Duc Chinh, Nguyen Thi Hanh

机构 * Hanoi University of Science and Technology(河内科学技术大学) Faculty of Interdisciplinary Digital Technology(跨学科数字技术学院) PHENIKAA University(PHENIKAA大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04058 2025-08-18 cs.CV 83%

LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding

Feng Xiao, Hongbin Xu, Guocan Zhao, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06293 2025-08-15 cs.CV 83%

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness

Qifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu, Bosheng Qin, Wenqiao Zhang, Yunfei Li, Juncheng Li, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09085 2025-08-13 cs.NI cs.AI cs.LG 83%

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring

Zihan Fang, Zheng Lin, Senkang Hu, Yihang Tao, Yiqin Deng, Xianhao Chen, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City and Department of Computer Science, City University of Hong Kong(香港JC STEM实验室及城市大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08589 2025-08-13 cs.CV 83%

DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding

Wenwen Yu, Zhibo Yang, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07803 2025-08-12 cs.CV 83%

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

Yushen Xu, Xiaosong Li, Zhenyu Kuang, Xiaoqi Cheng, Haishu Tan, Huafeng Li

机构 * School of Physics and Optoelectronic Engineering(物理与光电工程学院) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) School of Information Engineering and Automation(信息工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07681 2025-08-12 cs.LG cs.AI 83%

MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation

Yooseok Lim, ByoungJun Jeon, Seong-A Park, Jisoo Lee, Sae Won Choi, Chang Wook Jeong, Ho-Geol Ryu, Hongyeol Lee, Hyun-Lim Yang

机构 * Seoul National University Hospital(首尔国立大学医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06701 2025-08-12 cs.CV cs.AI cs.CL cs.LG cs.SD eess.AS 83%

MMFformer: Multimodal Fusion Transformer Network for Depression Detection

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆斯美国大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏