arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2512.02055 2025-12-03 cs.CV cs.AI 81%

Leveraging AI multimodal geospatial foundation models for improved near-real-time flood mapping at a global scale

利用AI多模态地理空间基础模型实现全球范围内的改进型实时洪水制图

Mirela G. Tulbure, Julio Caineta, Mark Broich, Mollie D. Gaines, Philippe Rufin, Leon-Friedrich Thomas, Hamed Alemohammad, Jan Hemmerling, Patrick Hostert

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究利用AI多模态地理空间基础模型提升全球实时洪水制图能力,通过微调TerraMind模型并对比不同配置,展示多模态数据整合在洪水检测中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01442 2025-12-02 cs.MM cs.AI 81%

PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis

PSA-MF:基于人格-情感对齐的多级融合用于多模态情感分析

Heng Xie, Kang Zhu, Zhengqi Wen, Jianhua Tao, Xuefei Liu, Ruibo Fu, Changsheng Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 PSA-MF通过引入人格-情感对齐和多级融合方法,提升多模态情感分析的识别性能。

Comments AAAI 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01214 2025-12-02 cs.CV cs.AI 81%

M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis

M4-BLIP:通过面部增强的局部分析推进多模态媒体篡改检测

Hang Wu, Ke Sun, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing(多媒体可信感知与高效计算重点实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 M4-BLIP通过引入面部增强的局部分析,提升多模态媒体篡改检测的准确性和可解释性。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08679 2025-12-02 cs.CV cs.AI 81%

MMIF-AMIN: Adaptive Loss-Driven Multi-Scale Invertible Dense Network for Multimodal Medical Image Fusion

MMIF-AMIN: 适应性损失驱动的多尺度可逆密集网络用于多模态医学图像融合

Tao Luo, Weihua Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 MMIF-AMIN通过可逆密集网络和多尺度互补特征提取模块,实现多模态医学图像融合的高效融合与精准诊断。

Comments This manuscript is withdrawn to allow for substantial expansion and restructuring. Based on recent research progress, we plan to add Generalization experiment and reorganize the manuscript structure to improve readability and logical flow. Thank you for your understanding and support

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01528 2025-12-01 cs.CL cs.AI q-bio.BM 81%

Leveraging Biomolecule and Natural Language through Multi-Modal Learning: A Survey

利用生物分子和自然语言通过多模态学习:一篇综述

Qizhi Pei, Zhimeng Zhou, Kaiyuan Gao, Jinhua Zhu, Yue Wang, Zun Wang, Tao Qin, Lijun Wu, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) Huazhong University of Science and Technology(华中科技大学) University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Shanghai AI Laboratory(上海人工智能实验室) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文综述了生物分子与自然语言多模态学习的最新进展,探讨了技术表示、多模态整合方法、应用实例及未来研究方向。

Comments 2025.11.28 Updated Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19887 2025-11-26 cs.CV cs.AI 81%

Distilling Cross-Modal Knowledge via Feature Disentanglement

通过特征解耦进行跨模态知识蒸馏

Junhong Liu, Yuan Zhang, Tao Huang, Wenchao Xu, Renyu Yang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出频率解耦的跨模态知识蒸馏方法,通过利用频域特征解耦和平衡跨模态知识转移,提升蒸馏效果。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18824 2025-11-25 cs.CV cs.CL 81%

Assessing the alignment between infants' visual and linguistic experience using multimodal language models

利用多模态语言模型评估婴儿的视觉和语言经验一致性

Alvin Wei Ming Tan, Jane Yang, Tarun Sepuri, Khai Loong Aw, Robert Z. Sparks, Zi Yin, Virginia A. Marchman, Michael C. Frank, Bria Long

机构 * Department of Psychology, Stanford University(心理学系,斯坦福大学) Department of Psychology, University of California, San Diego(心理学系,加州大学圣地亚哥分校) Department of Psychology, Tsinghua University(心理学系,清华大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 基于多模态语言模型评估婴儿视觉与语言经验一致性,揭示日常学习中视觉与语言对齐的稀有性及跨儿童差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17965 2025-11-25 cs.CV cs.MM 81%

Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification

信号:选择性交互与全局-局部对齐用于多模态对象重识别

Yangyang Liu, Yuhao Wang, Pingping Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.MM

AI总结 Signal通过选择性交互和全局-局部对齐框架提升多模态对象重识别的性能。

Comments Accepted by AAAI2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14169 2025-11-25 cs.CV cs.AI 81%

AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs

AdaTok: 一种基于对象感知表示的自适应令牌压缩方法,用于高效多模态大语言模型

Xinliang Zhang, Lei Zhu, Hangzhou He, Shuang Zeng, Ourui Fu, Jiakui Hu, Zhengjian Yao, Yanye Lu

机构 * Institute of Medical Technology, Peking University Health Science Center(北京大学医学部医学技术研究所) Department of Biomedical Engineering, Peking University(北京大学生物医学工程系) National Biomedical Imaging Center, Peking University(北京大学国家生物医学成像中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 AdaTok通过自适应令牌压缩方法,利用对象感知表示提升多模态大语言模型的效率,实现高压缩比与高性能的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14969 2025-11-20 eess.AS cs.AI cs.LG eess.IV eess.SP 81%

Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion

Zanxu Wang, Homayoon Beigi

机构 * Columbia University, New York, USA(哥伦比亚大学) Recognition Technologies, Inc., New York, USA(识别技术公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments 8 pages, 14 images, 3 tables, Recognition Technologies, Inc. Technical Report RTI-20251118-01

Journal ref Recognition Technologies, Inc. Technical Reports, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13794 2025-11-19 cs.CV cs.AI 81%

FusionFM: All-in-One Multi-Modal Image Fusion with Flow Matching

Huayi Zhu, Xiu Shu, Youqiang Xiong, Qiao Liu, Rui Chen, Di Yuan, Xiaojun Chang, Zhenyu He

机构 * Guangzhou Institute of Technology, Xidian University(广州理工大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03939 2025-11-07 cs.LG cs.AI cs.CL 81%

RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods

Raghav Sharma, Manan Mehta, Sai Tiger Raina

机构 * Northeastern University(东北大学) University of Southern California(南加州大学)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27208 2025-11-03 cs.CV cs.AI 81%

Multi-Modal Feature Fusion for Spatial Morphology Analysis of Traditional Villages via Hierarchical Graph Neural Networks

Jiaxin Zhang, Zehong Zhu, Junye Deng, Yunqin Li, and Bowen Wang

机构 * Architecture and Design College, Nanchang University(南昌大学建筑与设计学院) SANKEN, The University of Osaka(大阪大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09135 2025-10-30 cs.AI cs.CL cs.HC cs.LG 81%

Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation

Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K. Vail, Sunreeta Bhattacharya, Álvaro Fernández García, Kailana Baker-Matsuoka, Sheryl Mathew, Lori L. Holt, Fernando De la Torre

机构 * Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所) Department of Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校心理学系) Center for Perceptual Systems, The University of Texas at Austin(德克萨斯大学奥斯汀分校感知系统中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 22 pages, first three authors equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24777 2025-10-30 cs.CV cs.AI eess.IV 81%

Cross-Enhanced Multimodal Fusion of Eye-Tracking and Facial Features for Alzheimer's Disease Diagnosis

Yujie Nie, Jianzhang Ni, Yonglong Ye, Yuan-Ting Zhang, Yun Kwok Wing, Xiangqing Xu, Xin Ma, Lizhou Fan

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学) Engineering Research Center of Intelligent Unmanned System, Ministry of Education(智能无人机系统工程研究中心,教育部) Department of Psychiatry, The Chinese University of Hong Kong(心理学系,香港中文大学) Department of Electronic Engineering, The Chinese University of Hong Kong(电子工程系,香港中文大学) AICARE Lab, Guangdong Medical University(AICARE实验室,广东医科大学) Department of Neurology, Shandong University of Traditional Chinese Medicine Affiliated Hospital(神经内科,山东中医药大学附属医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 35 pages, 8 figures, and 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22507 2025-10-28 cs.CV cs.AI 81%

GateFuseNet: An Adaptive 3D Multimodal Neuroimaging Fusion Network for Parkinson's Disease Diagnosis

Rui Jin, Chen Chen, Yin Liu, Hongfu Sun, Min Zeng, Min Li, Yang Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The first two authors contributed equally to this work. Correspondence to: Yang Gao, E-mail: yang.gao@csu.edu.cn

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18279 2025-10-23 cs.CL cs.AI 81%

Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs

Yanhong Li, Zixuan Lan, Jiawei Zhou

机构 * Allen Institute for AI(艾伦人工智能研究所) University of Chicago(芝加哥大学) Stony Brook University(石溪大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings ("Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs")

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19257 2025-10-22 cs.CV cs.CL 81%

MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models

Yinan Xia, Yilei Jiang, Yingshui Tan, Xiaoyong Zhu, Xiangyu Yue, Bo Zheng

机构 * Future Lab, Alibaba Group(阿里巴巴集团未来实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00916 2025-10-21 cs.CV cs.AI 81%

Enhancing Osteoporosis Detection: An Explainable Multi-Modal Learning Framework with Feature Fusion and Variable Clustering

Mehdi Hosseini Chagahi, Saeed Mohammadi Dashtaki, Niloufar Delfan, Nadia Mohammadi, Farshid Rostami Pouria, Behzad Moshiri, Md. Jalil Piran, Oliver Faust

机构 * School of Electrical and Computer Engineering, College of Engineering, University of Tehran(塔里班大学电气与计算机工程学院) Department of Epidemiology, Shiraz University of Medical Science(谢尔兹医学科学大学流行病学系) Department of Computer Science and Engineering, Sejong University(世宗大学计算机科学与工程系) School of Computing and Information Science, Anglia Ruskin University(安格利亚 Ruskin 大学计算与信息科学学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13131 2025-10-16 cs.CV cs.MM 81%

OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment

Rongjun Chen, Chengsi Yao, Jinchang Ren, Xianxian Zeng, Peixian Wang, Jun Yuan, Jiawen Li, Huimin Zhao, Xu Lu

机构 * School of Computer Science, Guangdong Polytechnic Normal University(广东 polytechnic 正规大学计算机学院)

专题命中 多模态训练与对齐 :image-text(title);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05839 2025-10-15 cs.MM cs.CV 81%

Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality

Hengyang Zhou, Yiwei Wei, Jian Yang, Zhenyu Zhang

机构 * Nanjing University(南京大学) China University of Petroleum(中国石油大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10406 2025-10-14 cs.CV cs.AI cs.LG 81%

Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes

Zhao-Yang Wang, Jieneng Chen, Jiang Liu, Yuxiang Guo, Rama Chellappa

机构 * Johns Hopkins University(约翰霍普金斯大学) Advanced Micro Devices, Inc.(先进微器件公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.06145 2025-10-13 cs.CV cs.AI cs.HC cs.LG stat.ML 81%

Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training

Mahdi Abavisani, Hamid Reza Vaezi Joze, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 1165-1174

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06498 2025-10-13 cs.LG cs.AI cs.CV stat.ML 81%

Deep Multimodal Subspace Clustering Networks

Mahdi Abavisani, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref IEEE Journal of Selected Topics in Signal Processing, vol. 12, no. 6, pp. 1601-1614, Dec. 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07513 2025-10-10 cs.LG cs.AI cs.CV cs.DB 81%

MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis

Qinghua Liu, Sam Heshmati, Zheda Mai, Zubin Abraham, John Paparrizos, Liu Ren

机构 * The Ohio State University(俄亥俄州立大学) Bosch Research North America(博世北美研究)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24776 2025-09-30 cs.CV cs.AI 81%

VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding

Yizhuo Ding, Mingkang Chen, Zhibang Feng, Tong Xiao, Wanying Qu, Wenqi Shao, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shenzhen University(深圳大学) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24734 2025-09-30 cs.LG cs.AI cs.CV 81%

A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity

Giordano Cicchetti, Eleonora Grassucci, Danilo Comminiello

机构 * Department of Information Engineering, Electronics, and Telecommunications(信息工程、电子与电信系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23109 2025-09-30 cs.AI cs.CV 81%

AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors

Junyang Zhang, Tianyi Zhu, Thierry Tambe

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 31 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22729 2025-09-30 cs.CL cs.AI 81%

Multi-Modal Sentiment Analysis with Dynamic Attention Fusion

Sadia Abdulhalim, Muaz Albaghdadi, Moshiur Farazi

机构 * University of Doha for Science and Technology(多哈科学技术大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CL、cs.AI

Comments Paper accepted for presentation at the ACS/IEEE 22nd International Conference on Computer Systems and Applications (AICCSA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22261 2025-09-29 cs.AI cs.CL 81%

InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning

Guanghao Zhu, Zhitian Hou, Zeyu Liu, Zhijie Sang, Congkai Xie, Hongxia Yang

机构 * The Hong Kong Polytechnic University(香港理工大学) Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏