arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2507.09747 2025-10-07 cs.NE 82%

BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings

Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei, Quanying Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25278 2025-10-01 cs.LG stat.ML 82%

MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time Series

Payal Mohapatra, Yueyuan Sui, Akash Pandey, Stephen Xia, Qi Zhu

机构 * Northwestern University(西北大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted to Neurips 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19018 2025-09-24 cs.LG 82%

OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment

Teng Xiao, Zuchao Li, Lefei Zhang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09114 2025-09-12 cs.IR 82%

Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation

Kelin Ren, Chan-Yang Ju, Dong-Ho Lee

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02943 2025-09-04 cs.IR 82%

Knowledge graph-based personalized multimodal recommendation fusion framework

Yu Fang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20279 2025-08-29 cs.CV cs.AI cs.CL 82%

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

Zhuoran Yu, Yong Jae Lee

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16147 2025-08-25 cs.IR 82%

Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction

Ao Zhou, Mingsheng Tu, Luping Wang, Tenghao Sun, Zifeng Cheng, Yafeng Yin, Zhiwei Jiang, Qing Gu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13971 2025-08-20 eess.AS cs.CL cs.HC cs.LG cs.MM 82%

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience

Andrew Chang, Chenkai Hu, Ji Qi, Zhuojian Wei, Kexin Zhang, Viswadruth Akkaraju, David Poeppel, Dustin Freeman

机构 * New York UniversityUSA(纽约大学) Max Planck SocietyGermany(马克斯·普朗克研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.MM、eess.AS

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02133 2025-08-19 cs.HC 82%

Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs

Yitong Zhu, Lei Han, Guanxuan Jiang, PengYuan Zhou, Yuyang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11886 2025-08-19 cs.CV cs.AI cs.CL cs.LG eess.IV 82%

EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models

Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Shao Tang, Sayan Ghosh, Xuanzhao Dong, Rajat Koner, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) LinkedIn Corporation(领英公司) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05476 2025-08-08 eess.IV 82%

MM2CT: MR-to-CT translation for multi-modal image fusion with mamba

Chaohui Gong, Zhiying Wu, Zisheng Huang, Gaofeng Meng, Zhen Lei, Hongbin Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01805 2025-08-05 cs.NI 82%

M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks

Yongjie Zeng, Hongyang Du

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00926 2025-08-05 cs.LG 82%

Hybrid Hypergraph Networks for Multimodal Sequence Data Classification

Feng Xu, Hui Wang, Yuting Huang, Danwei Zhang, Zizhu Fan

机构 * Feng Xu 1,2(作者1单位) Hui Wang 1(作者1单位) Yuting Huang 3(作者3单位) Danwei Zhang 4(作者4单位) Zizhu Fan 5(作者5单位)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03248 2025-07-30 cs.CV cs.AI cs.CL 82%

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Yiwu Zhong, Zhuoming Liu, Yin Li, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19628 2025-07-28 cs.CV cs.CL cs.LG cs.MM 82%

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Qiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye, Zhanpeng Zen, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22334 2025-07-24 cs.CL cs.AI cs.CV cs.LG 82%

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Lai Wei, Yuting Li, Kaipeng Zheng, Chen Wang, Yue Wang, Linghe Kong, Lichao Sun, Weiran Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23115 2025-07-01 cs.CV cs.AI cs.CL 82%

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Haonan Chen, Hong Liu, Yuping Luo, Liang Wang, Nan Yang, Furu Wei, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Stanford University(斯坦福大学) Microsoft Corporation(微软公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Homepage: https://haon-chen.github.io/MoCa/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18023 2025-06-26 cs.CV cs.AI cs.CL 82%

PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding

Kui Huang, Xinrong Chen, Wenyu Lv, Jincheng Liao, Guanzhong Wang, Yi Liu

机构 * Baidu Inc.(百度公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21755 2025-06-24 cs.CV cs.AI cs.CL cs.LG 82%

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering

Chengyue Huang, Brisa Maneechotesuwan, Shivang Chopra, Zsolt Kira

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16744 2025-06-23 cs.LG cs.RO eess.SP 82%

IsoNet: Causal Analysis of Multimodal Transformers for Neuromuscular Gesture Classification

Eion Tyacke, Kunal Gupta, Jay Patel, Rui Li

机构 * Dept. of Electrical and Computer Engineering(电气与计算机工程系) New York University(纽约大学) Tandon School of Engineering(坦顿工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10452 2025-06-13 cs.CV cs.CL cs.LG cs.MM 82%

Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts

Guowei Zhong, Ruohong Huan, Mingzhen Wu, Ronghua Liang, Peng Chen

机构 * College of Computer Science and Technology, Zhejiang University of Technology(浙江工业大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Submitted to TAC. The code is available at https://github.com/gw-zhong/CIDer

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01067 2025-06-12 cs.AI cs.CL cs.CV cs.HC cs.LG 82%

Human-like object concept representations emerge naturally in multimodal large language models

Changde Du, Kaicheng Fu, Bincheng Wen, Yi Sun, Jie Peng, Wei Wei, Ying Gao, Shengpei Wang, Chuncheng Zhang, Jinpeng Li, Shuang Qiu, Le Chang, Huiguang He

机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑启发智能技术重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Neuroscience, State Key Laboratory of Brain Cognition and Brain-Inspired Intelligence Technology(神经科学研究所) CAS Center for Excellence in Brain Science and Intelligence Technology(中国科学院脑科学与智能技术卓越创新中心) Chinese Academy of Sciences(中国科学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published on Nature Machine Intelligence

Journal ref Nature Machine Intelligence, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18956 2025-06-11 cs.CV cs.AI cs.LG cs.MM 82%

How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation

Yining Pan, Qiongjie Cui, Xulei Yang, Na Zhao

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学) Institute for Infocomm Research (I2R), A*STAR, Singapore(信息与通信研究院(I2R))

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted at the 2025 International Conference on Machine Learning (ICML)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20188 2025-05-27 cs.LG cs.IR 82%

Research on feature fusion and multimodal patent text based on graph attention network

Zhenzhen Song, Ziwei Liu, Hongji Li

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18536 2025-05-27 cs.CL cs.AI cs.CV 82%

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Haoyuan Sun, Jiaqi Wu, Bo Xia, Yifu Luo, Yifei Zhao, Kai Qin, Xufei Lv, Tiantian Zhang, Yongzhe Chang, Xueqian Wang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08334 2025-05-22 cs.CV cs.AI cs.IR cs.MM 82%

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval

Yeong-Joon Ju, Ho-Joong Kim, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23333 2025-04-01 cs.IR cs.AI cs.CL cs.CV 82%

Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics

Jing Zhu, Mingxuan Ju, Yozen Liu, Danai Koutra, Neil Shah, Tong Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21775 2025-03-31 cs.CV cs.AI cs.CL cs.GR cs.LG 82%

StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion

Ziyu Guo, Young Yoon Lee, Joseph Liu, Yizhak Ben-Shabat, Victor Zordan, Mubbasir Kapadia

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://stylemotif.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20633 2025-03-27 cs.LG 82%

Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning

Sashuai Zhou, Hai Huang, Yan Xia

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

Comments 6 pages,3 figures

Journal ref ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16023 2025-03-21 cs.CR 82%

BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models

Zenghui Yuan, Jiawen Shi, Pan Zhou, Neil Zhenqiang Gong, Lichao Sun

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract)

Comments This paper is accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏