arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2505.17436 2025-05-26 cs.AI 79%

Scaling Up Biomedical Vision-Language Models: Fine-Tuning, Instruction Tuning, and Multi-Modal Learning

Cheng Peng, Kai Zhang, Mengxian Lyu, Hongfang Liu, Lichao Sun, Yonghui Wu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12453 2025-05-26 cs.MM 79%

Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding

Hanlei Zhang, Qianrui Zhou, Hua Xu, Jianhua Su, Roberto Evans, Kai Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16703 2025-05-23 cs.CL 79%

Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs

Zeping Yu, Sophia Ananiadou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18047 2025-05-23 cs.CV cs.LG 79%

Progressive Local Alignment for Medical Multimodal Pre-training

Huimin Yan, Xian Yang, Liang Bai, Jiye Liang

专题命中 多模态训练与对齐 :multimodal(title);image-text(abstract);分类 cs.CV

Comments We are currently revising the methodology described in the manuscript to improve its clarity. We have decided to withdraw the current version until a more robust and complete version is ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14715 2025-05-22 eess.IV cs.CV 79%

A Comprehensive Review of Techniques, Algorithms, Advancements, Challenges, and Clinical Applications of Multi-modal Medical Image Fusion for Improved Diagnosis

Muhammad Zubair, Muzammil Hussai, Mousa Ahmad Al-Bashrawi, Malika Bendechache, Muhammad Owais

机构 * Interdisciplinary Research Center for Finance and Digital Economy, King Fahd University of Petroleum and Minerals(金融与数字经济交叉研究中心,国王法赫德石油和矿物大学) Department of Software Engineering, Faculty of Information Technology, Al-Ahliyya Amman University(软件工程系,信息科技学院,阿尔阿赫利亚大学) Department of Information Systems and Operations Management, King Fahd University of Petroleum and Minerals(信息系统与运营管理系,国王法赫德石油和矿物大学) ADAPT Research Centre, School of Computer Science, University of Galway(ADAPT研究中心,计算机科学学院,Galway大学) Department of Mechanical and Nuclear Engineering, Khalifa University(机械与核工程系,哈利法大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments computerized medical imaging and graphics Journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11997 2025-05-21 cs.CV 79%

Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance

Mingcheng Qu, Guang Yang, Donglin Di, Tonghua Su, Yue Gao, Yang Song, Lei Fan

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Software, Tsinghua University(清华大学软件学院) School of Computer Science and Engineering, UNSW Sydney(新南威尔士大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments accepted by IJCAI2025 Code: https://github.com/MCPathology/MRePath

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09018 2025-05-21 cs.CV cs.LG 79%

Multimodal Fusion of Glucose Monitoring and Food Imagery for Caloric Content Prediction

Adarsh Kumar

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments The manuscript was submitted without proper consideration of institutional policies. Upon review with professor, it was found that the content is subject to licensing restrictions which prohibit public dissemination in its current form. Therefore, I am withdrawing the paper to comply with these requirements

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13175 2025-05-20 cs.AI 79%

Enhancing LLMs for Time Series Forecasting via Structure-Guided Cross-Modal Alignment

Siming Sun, Kai Zhang, Xuejun Jiang, Wenchao Meng, Qinmin Yang

机构 * Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12782 2025-05-20 cs.GR cs.CV cs.IR cs.IT math.IT 79%

AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning

Kai Zhang, Xingyu Chen, Xiaofeng Zhang

机构 * Kai Zhang 1(某机构) Xingyu Chen 3(某机构) Xiaofeng Zhang 2(某机构)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12744 2025-05-20 cs.AI 79%

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation

Weiliang Tang, Dong Jing, Jia-Hui Pan, Zhiwu Lu, Yun-Hui Liu, Li Erran Li, Mingyu Ding, Chi-Wing Fu

机构 * Department of Computer Science(计算机科学系) The Chinese University of Hong Kong(香港中文大学) Gaoling School of Artificial Intelligence(九龙人工智能学院) Renmin University of China(中国人民大学) AWS AI Amazon(AWS人工智能亚马逊) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 17 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09897 2025-05-20 cs.CV 79%

TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models

Jaewoo Lee, Keyang Xuan, Chanakya Ekbote, Sandeep Polisetty, Yi R. Fung, Paul Pu Liang

机构 * University of North Carolina Chapel Hill(北卡罗来纳大学教堂山分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Massachusetts Institute of Technology(麻省理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07372 2025-05-20 eess.IV cs.CV 79%

Multi-modal MRI Translation via Evidential Regression and Distribution Calibration

Jiyao Liu, Shangqi Gao, Yuxin Li, Lihao Liu, Xin Gao, Zhaohu Xing, Junzhi Ning, Yanzhou Su, Xiao-Yong Zhang, Junjun He, Ningsheng Xu, Xiahai Zhuang

机构 * Fudan University(复旦大学) University of Cambridge(剑桥大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Fuzhou University(福州市大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Early accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10862 2025-05-19 cs.CL 79%

Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?

Tairan Fu, Miguel González, Javier Conde, Elena Merino-Gómez, Pedro Reviriego

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 6 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06592 2025-05-13 cs.CV 79%

Batch Augmentation with Unimodal Fine-tuning for Multimodal Learning

H M Dipu Kabir, Subrota Kumar Mondal, Mohammad Ali Moni

机构 * AI and Cyber Futures Institute, Charles Sturt University, Australia(人工智能与网络未来研究所,查尔斯·斯特劳特大学,澳大利亚) Rural Health Research Institute, Charles Sturt University, Australia(农村健康研究研究所,查尔斯·斯特劳特大学,澳大利亚) School of Computer Science and Engineering, Macau University of Science and Technology, Macao(计算机科学与工程学院,澳门科学技术大学,澳门)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06333 2025-05-13 cs.LG cs.AI 79%

NSF-MAP: Neurosymbolic Multimodal Fusion for Robust and Interpretable Anomaly Prediction in Assembly Pipelines

Chathurangi Shyalika, Renjith Prasad, Fadi El Kalach, Revathy Venkataramanan, Ramtin Zand, Ramy Harik, Amit Sheth

机构 * Artificial Intelligence Institute, University of South Carolina(南卡罗来纳大学人工智能研究所) Clemson Composites Center, Clemson University(克莱姆森大学复合材料中心) Intelligent Circuits, Architectures and Systems Lab, University of South Carolina(南卡罗来纳大学智能电路、架构和系统实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 7 figures, 2 tables, IJCAI 2025 (International Joint Conferences on Artificial Intelligence) Special Track on AI4Tech: AI Enabling Critical Technologies

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05877 2025-05-13 cs.LG cs.AI 79%

Multi-Modal Molecular Representation Learning via Structure Awareness

Rong Yin, Ruyue Liu, Xiaoshuai Hao, Xingrui Zhou, Yong Liu, Can Ma, Weiping Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyberspace Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Xidian University(西安电子科技大学) Renmin University of China(中国人民大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by IEEE Transactions on Image Processing (TIP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20698 2025-05-12 cs.CV cs.IR 79%

MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion

Saron Samuel, Dan DeGenaro, Jimena Guallar-Blasco, Kate Sanders, Oluwaseun Eisape, Tanner Spendlove, Arun Reddy, Alexander Martin, Andrew Yates, Eugene Yang, Cameron Carpenter, David Etter, Efsun Kayi, Matthew Wiesner, Kenton Murray, Reno Kriz

机构 * Stanford University(斯坦福大学) Georgetown University(乔治城大学) Johns Hopkins University(约翰霍普金斯大学) UC Berkeley(伯克利大学) BYU

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07503 2025-05-09 cs.AI cs.LG 79%

Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems

Ibrahim Alabdulmohsin, Xiaohua Zhai

机构 * Google Deepmind(谷歌DeepMind)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22367 2025-05-07 q-bio.QM cs.AI cs.LG 79%

MAMMAL -- Molecular Aligned Multi-Modal Architecture and Language

Yoel Shoshan, Moshiko Raboh, Michal Ozery-Flato, Vadim Ratner, Alex Golts, Jeffrey K. Weber, Ella Barkan, Simona Rabinovici-Cohen, Sagi Polaczek, Ido Amos, Ben Shapira, Liam Hazan, Matan Ninio, Sivan Ravid, Michael M. Danziger, Yosi Shamay, Sharon Kurant, Joseph A. Morrone, Parthasarathy Suryanarayanan, Michal Rosen-Zvi, Efrat Hexter

机构 * IBM Research-Israel(IBM研究以色列分公司) IBM TJ Watson Research Center(IBM TJ Watson研究中心) Faculty of Biomedical Engineering(生物医学工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02529 2025-05-06 eess.IV cs.CV 79%

RobSurv: Vector Quantization-Based Multi-Modal Learning for Robust Cancer Survival Prediction

Aiman Farooq, Azad Singh, Deepak Mishra, Santanu Chaudhury

机构 * Indian Institute of Technology Jodhpur(印度理工学院朱道尔分校) Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02486 2025-05-06 cs.LG cs.AI 79%

SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning

Jinpeng Chen, Runmin Cong, Yuzhi Zhao, Hongzheng Yang, Guangneng Hu, Horace Ho Shing Ip, Sam Kwong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09256 2025-05-06 cs.SD eess.AS 79%

Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation

Yifei Xin, Zhihong Zhu, Xuxin Cheng, Xusheng Yang, Yuexian Zou

机构 * YifeiXin(西菲·欣) ZhihongZhu(朱之宏) XuxinCheng(程旭鑫) XushengYang(杨旭生) YuexianZou(邹岳先)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 eess.AS

Comments Accepted by Interspeech2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01766 2025-05-06 cs.CV cs.RO 79%

Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement

Long Bai, Boyi Ma, Ruohan Wang, Guankun Wang, Beilei Cui, Zhongliang Jiang, Mobarakol Islam, Zhe Min, Jiewen Lai, Nassir Navab, Hongliang Ren

机构 * Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系) Chair for Computer Aided Medical Procedures, Technical University of Munich(慕尼黑技术大学计算机辅助医疗程序主席职位) Department of Biomedical Engineering, University of Toronto(多伦多大学生物医学工程系) Center for Computational and Molecular Biology, Brown University(布朗大学计算与分子生物学中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21281 2025-05-01 cs.CV 79%

Mamba Based Feature Extraction And Adaptive Multilevel Feature Fusion For 3D Tumor Segmentation From Multi-modal Medical Image

Zexin Ji, Beiji Zou, Xiaoyan Kui, Hua Li, Pierre Vera, Su Ruan

机构 * School of Computer Science and Engineering, Central South University, Changsha, 410083, China(计算机科学与工程学院,中南大学,长沙,410083,中国) Hunan Engineering Research Center of Machine Vision and Intelligent Medicine, Central South University, Changsha, 410083, China(机器视觉与智能医学工程研究中心,中南大学,长沙,410083,中国) Department of Radiation Oncology, Washington University in St. Louis, USA(放射肿瘤科,华盛顿大学圣路易斯分校,美国) Department of Nuclear Medicine, Henri Becquerel Cancer Center, Rouen, France(核医学科,亨利·贝克勒尔癌症中心,鲁昂,法国) University of Rouen-Normandy, AMIS - QuantIF UR 4108, F-76000, Rouen, France(鲁昂-诺曼底大学,AMIS - QuantIF UR 4108,法国,F-76000,鲁昂,法国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20362 2025-04-30 cs.CV 79%

TTTFusion: A Test-Time Training-Based Strategy for Multimodal Medical Image Fusion in Surgical Robots

Qinhua Xie, Hao Tang

机构 * School of Data Science and Engineering, East China Normal University(东华大学数据科学与工程学院) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20178 2025-04-30 cs.CV cs.LG 79%

A Transformer-based Multimodal Fusion Model for Efficient Crowd Counting Using Visual and Wireless Signals

Zhe Cui, Yuli Li, Le-Nam Tran

机构 * School of Electrical and Electronic Engineering, University College Dublin(都柏林大学电子与电气工程学院) College of Electrical Engineering and Automation, Shandong University of Science and Technology(山东科技大学电气工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments This paper was accepted at IEEE WCNC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19839 2025-04-29 cs.CV 79%

SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation

Yulong Guo, Zilun Zhang, Yongheng Shang, Tiancheng Zhao, Shuiguang Deng, Yingchun Yang, Jianwei Yin

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Haina Institute of Zhejiang University, Advanced Technology Institute, Zhejiang University(浙江大学海纳研究院、先进技术研究院) Binjiang Research Institute of Zhejiang University(浙江大学滨江研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments None

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18961 2025-04-29 cs.IR cs.AI 79%

Feature Fusion Revisited: Multimodal CTR Prediction for MMCTR Challenge

Junjie Zhou

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家实验室,南京大学,中国) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments A technical report for the MMCTR Challenge held by EReL@MIR Workshop at WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17223 2025-04-25 cs.CV 79%

Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion

Mengyu Qiao, Runze Tian, Yang Wang

机构 * North China University of Technology(华北理工大学) Ultramain Systems, Inc.(Ultramain系统公司)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15848 2025-04-23 cs.CL 79%

Exploring Cognitive and Aesthetic Causality for Multimodal Aspect-Based Sentiment Analysis

Luwei Xiao, Rui Mao, Shuai Zhao, Qika Lin, Yanhao Jia, Liang He, Erik Cambria

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学 Saw Swee Hock 公共卫生学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by TAFFC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏