arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-03 至 2026-02-03 共收录 24 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 24 篇

2502.05568 2026-02-03 cs.CL cs.AI cs.LG 84%

Large Multimodal Models for Low-Resource Languages: A Survey

大规模多模态模型在低资源语言中的应用:综述

Marian Lupascu, Ana-Cristina Rogoz, Mihai Sorin Stupariu, Radu Tudor Ionescu

机构 * Department of Computer Science, University of Bucharest(布加勒斯特大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文综述了大规模多模态模型在低资源语言中的应用,分析了适应技术及挑战,提供了方法导向的贡献比较和开放源代码资源。

Comments Accepted in Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00505 2026-02-03 cs.CV cs.AI 84%

Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language Models

稀疏快捷键:促进多模态大语言模型中的高效融合

Jingrui Zhang, Feng Liang, Yong Zhang, Wei Wang, Runhao Zeng, Xiping Hu

机构 * Guangdong-Hong Kong-Macao Joint Laboratory for Emotion Intelligence and Pervasive Computing, Artificial Intelligence Research Institute, Shenzhen MSU-BIT University, Shenzhen(粤港澳大湾区情感智能与 pervasive 计算联合实验室,人工智能研究院,深圳MSU-BIT大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 SparseCut通过稀疏快捷连接和多粒度特征融合模块,提升多模态大语言模型在跨模态融合中的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00143 2026-02-03 cs.LG cs.AI cs.CV 84%

Invariant Representation Guided Multimodal Sentiment Decoding with Sequential Variation Regularization

不变表示引导的多模态情感解码与序列变化正则化

Guoyang Xu, Zhenxi Song, Junqi Xue, Yuxin Liu, Zirui Wang, Zhiguo Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学深圳校区)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种通过模态不变融合和序列变化正则化相结合的方法,以提升多模态情感解码的稳定性和准确性。

Comments change Title, Authors, Abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20198 2026-02-03 cs.CV 84%

A Survey of Token Compression for Efficient Multimodal Large Language Models

多模态大语言模型高效性中的标记压缩综述

Kele Shao, Keda Tao, Kejia Zhang, Sicheng Feng, Mu Cai, Yuzhang Shang, Haoxuan You, Can Qin, Yang Sui, Huan Wang

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) Xiamen University(厦门大学) National University of Singapore(新加坡国立大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Central Florida(佛罗里达大学) Salesforce AI Research(Salesforce AI研究) Rice University(德克萨斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文综述了多模态大语言模型中标记压缩技术,分类讨论了图像、视频和音频三种模态的压缩方法及其机制,旨在推动该领域的发展。

Comments For ongoing updates and to track the latest advances in this promising area, we maintain a public repository: https://github.com/cokeshao/Awesome-Multimodal-Token-Compression

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00914 2026-02-03 cs.CL cs.AI cs.CY cs.SD eess.AS 82%

A Baseline Multimodal Approach to Emotion Recognition in Conversations

一种用于对话中情感识别的基线多模态方法

Víctor Yeste, Rodrigo Rivas-Arévalo

机构 * School of Science, Engineering and Design, Universidad Europea de Valencia(科学、工程与设计学院,欧洲大学 Valencia)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本文提出了一种基于Transformer文本分类器和自监督语音模型的多模态基线方法,用于对话中情感识别,并通过实验展示了多模态融合的优势。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15388 2026-02-03 cs.CV cs.AI cs.CL 82%

LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models

LLaVA-PruMerge: 适应性令牌减少用于高效的大多模态模型

Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, Yan Yan

机构 * UCF(佛罗里达大学) UW-Madison(威斯康星大学麦迪逊分校) USC(南加州大学) UIC(伊利诺伊大学香槟分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 LLaVA-PruMerge通过自适应视觉令牌减少策略,显著降低视觉令牌数量而不影响性能,适用于高效的大多模态模型。

Comments Accepted to ICCV 2025. First Version is released in 2024/03

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04356 2026-02-03 cs.RO cs.AI cs.CV 81%

UNIC: Learning Unified Multimodal Extrinsic Contact Estimation

UNIC: 学习统一的多模态外在接触估计

Zhengtong Xu, Yuki Shirai

机构 * Edwardson School of Industrial Engineering, Purdue University(帕克大学工业工程学院) Mitsubishi Electric Research Laboratories(三菱电机研究实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 UNIC通过统一多模态框架实现无需先验知识的外在接触估计,提升手部操作的鲁棒性和适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00347 2026-02-03 cs.CV cs.AI 81%

AdaFuse: Adaptive Multimodal Fusion for Lung Cancer Risk Prediction via Reinforcement Learning

AdaFuse:基于强化学习的自适应多模态融合用于肺癌风险预测

Chongyu Qu, Zhengyi Lu, Yuxiang Lai, Thomas Z. Li, Junchao Zhu, Junlin Guo, Juming Xiong, Yanfan Zhu, Yuechen Yang, Allen J. Luna, Kim L. Sandler, Bennett A. Landman, Yuankai Huo

机构 * Vanderbilt University(范德比尔特大学) Vanderbilt University Medical Center(范德比尔特大学医学中心) Emory University(埃默里大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 AdaFuse利用强化学习实现自适应多模态融合,通过动态选择和融合模态提升肺癌风险预测的AUC性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01954 2026-02-03 cs.CV 79%

Beyond Open Vocabulary: Multimodal Prompting for Object Detection in Remote Sensing Images

超越开放词汇:面向遥感图像的目标检测多模态提示

Shuai Yang, Ziyue Huang, Jiaxin Chen, Qingjie Liu, Yunhong Wang

机构 * School of Computer Science and Engineering, Beihang University, Beijing, China(计算机科学与工程学院,北京航空航天大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 RS-MPOD提出了一种多模态开放词汇检测框架,通过整合视觉和文本提示提升遥感图像中目标检测的稳定性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01833 2026-02-03 cs.MM 79%

Mixture of Disentangled Experts with Missing Modalities for Robust Multimodal Sentiment Analysis

多模态情感分析中缺失模态的解耦专家混合方法

Xiang Li, Xiaoming Zhang, Dezhuang Miao, Xianfu Cheng, Dawei Li, Honggui Han, Zhoujun Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 DERL通过解耦专家和多层次重建策略,提升多模态情感分析在缺失模态下的鲁棒性和表现

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00946 2026-02-03 cs.CV 79%

ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models

ConsensusDrop:融合视觉与跨模态显著性以提高视觉语言模型的效率

Dhruv Parikh, Haoyang Fan, Rajgopal Kannan, Viktor Prasanna

机构 * University of Southern California(南加州大学) DEVCOM Army Research Office(陆军研究办公室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 ConsensusDrop通过融合视觉与跨模态显著性,提高视觉语言模型的效率和准确性,优于现有token修剪方法。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00730 2026-02-03 cs.IR 78%

Towards Trustworthy Multimodal Recommendation

迈向可信的多模态推荐

Zixuan Li

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种多模态推荐系统中的可信度校正方法,通过学习模态特征的软对应关系以提升鲁棒性,并验证了交互级可信度的两个关键观察。

Comments Preprint, 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02399 2026-02-03 cs.CR 78%

AtomGraph: Tackling Atomicity Violation in Smart Contracts using Multimodal GCNs

AtomGraph: 通过多模态GCN解决智能合约中的原子性违规问题

Xiaoqi Li, Zongwei Li, Wenkai Li, Zeng Zhang, Lei Xie

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 AtomGraph通过多模态GCN检测智能合约中的原子性违规,实现高准确率和F1分数,优于现有工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01832 2026-02-03 cs.AI 70%

Synesthesia of Vehicles: Tactile Data Synthesis from Visual Inputs

车辆的联觉:从视觉输入合成触觉数据

Rui Wang, Yaoguang Cao, Yuyi Chen, Jianyi Xu, Zhuoyang Li, Jiachen Shang, Shichun Yang

机构 * Dept. of Transportation Science, Beihang Univ.(北京航空航天大学交通运输科学系) State Key Lab of Intelligent Transportation System, Beihang Univ.(北京航空航天大学智能交通系统国家重点实验室) Hangzhou International Innovation Institute, Beihang Univ.(杭州国际创新研究院) China Software Testing Center(Ministry of Industry and Information Technology Software and Integrated Circuit Promotion Center)(中国软件测试中心(工业和信息化部软件与集成电路促进中心))

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出基于视觉输入的触觉数据合成方法,通过跨模态时空对齐和潜在扩散模型提升自动驾驶车辆的触觉感知能力,从而增强安全性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01610 2026-02-03 cs.AI cs.LG 70%

ToPT: Task-Oriented Prompt Tuning for Urban Region Representation Learning

为城市区域表示学习设计的任务导向提示微调:ToPT

Zitao Guo, Changyang Jiang, Tianhong Zhao, Jinzhou Cao, Genan Dai, Bowen Zhang

机构 * College of Applied Science, Shenzhen University, Shenzhen, China(深圳大学应用科学学院) School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China(深圳科技大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 ToPT通过空间一致融合和任务对齐提升城市区域表示学习,实现任务导向的提示微调,提升多个城市任务的性能。

Comments The paper has been accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11512 2026-02-03 cs.RO cs.CV 70%

Collaborative Representation Learning for Alignment of Tactile, Language, and Vision Modalities

协同表征学习用于触觉、语言和视觉模态的对齐

Yiyun Zhou, Mingjing Xu, Jingwei Shi, Quanjiang Li, Jingyuan Chen

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 TLV-CoRe通过协同表征学习提升触觉、语言和视觉模态的跨模态对齐与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07593 2026-02-03 cs.RO cs.AI cs.CV cs.SY eess.IV eess.SY 62%

Vision-Proprioception Fusion with Mamba2 in End-to-End Reinforcement Learning for Motion Control

基于Mamba2的视觉-本体感知融合在端到端强化学习中的运动控制

Xiaowen Tao, Yinuo Wang, Jinzhao Zhou

机构 * School of Computer Science and Statistics, Trinity College Dublin(都柏林三一学院计算机科学与统计学系) Faculty of Engineering and Information Technology, University of Technology Sydney(新南威尔士大学理工学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于SSD-Mamba2的视觉-本体感知融合框架,通过端到端强化学习提升运动控制的效率和安全性。

Comments 6 figures and 8 tables. This paper has been accepted by Advanced Engineering Informatics

Journal ref Advanced Engineering Informatics, vol. 71, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01593 2026-02-03 cs.CV 57%

Samba+: General and Accurate Salient Object Detection via A More Unified Mamba-based Framework

Samba+: 通过更统一的Mamba框架实现通用且准确的显著目标检测

Wenzhuo Zhao, Keren Fu, Jiahao He, Xiaohong Liu, Qijun Zhao, Guangtao Zhai

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) National Key Laboratory of Fundamental Science on Synthetic Vision, Sichuan University(合成视觉基础科学国家级重点实验室) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院) School of Information Science and Electronic Engineering, Shanghai Jiao Tong University(上海交通大学信息科学与电子工程学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 Samba+通过统一的Mamba框架实现多模态显著目标检测,提升模型通用性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01522 2026-02-03 cs.LG cs.CV 57%

When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning

何时秩-1足够?基于几何的初始化方法用于参数高效微调

Haoran Zhao, Soyeon Caren Han, Eduard Hovy

机构 * School of Computing and Information Systems, University of Melbourne, Melbourne, Australia(计算机与信息系统学院,墨尔本大学,墨尔本,澳大利亚)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出Gap-Init方法,通过几何感知初始化稳定秩-1 LoRA微调,证明初始对齐对训练稳定性的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16600 2026-02-03 cs.AI 57%

You Only Forward Once: An Efficient Compositional Judging Paradigm

你只需一次转发:一种高效的组合判断范式

Tianlong Zhang, Hongwei Xue, Shilin Yan, Di Wu, Chen Xu, Guannan Zhang, Yunyun Yang

机构 * School of Science, Harbin Institute of Technology, Shenzhen, Guangdong Province, China(哈尔滨工业大学深圳校区科学学院) Accio, Alibaba Group, Hangzhou, Zhejiang Province, China(阿里集团杭州Accio)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 YOFO通过单次前向传递高效判断多模态大语言模型的结构化要求,实现速度与可解释性的平衡,同时支持依赖分析和事后CoT。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00574 2026-02-03 cs.AI 57%

Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings

通过潜在嵌入学习模态混合的思考链推理

Yifei Shao, Kun Zhou, Ziming Xu, Mohammad Atif Quamar, Shibo Hao, Zhen Wang, Zhiting Hu, Biwei Huang

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文提出模态混合CoT方法,通过潜在嵌入结合视觉与文本信息,提升多模态推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00561 2026-02-03 cs.AI 57%

Uncovering Latent Communication Patterns in Brain Networks via Adaptive Flow Routing

通过自适应流路由揭示脑网络中的潜在通信模式

Tianhao Huang, Guanghui Min, Zhenyu Lei, Aiying Zhang, Chen Chen

机构 * University of Virginia(弗吉尼亚大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出AFR-Net,通过神经通信动态视角融合SC和FC,揭示潜在神经通路,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01916 2026-02-03 cs.RO 50%

ForSim: Stepwise Forward Simulation for Traffic Policy Fine-Tuning

ForSim:基于逐步前向模拟的交通策略微调

Keyu Chen, Wenchao Sun, Hao Cheng, Zheng Fu, Sifa Zheng

机构 * School of Vehicle and Mobility, Tsinghua University(车辆与移动系统学院,清华大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 ForSim通过逐步闭环前向模拟范式提升交通模拟的保真度,结合组相对优化微调交通策略,提高安全性与效率。

Comments Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00974 2026-02-03 cs.LG 50%

Forest-Guided Semantic Transport for Label-Supervised Manifold Alignment

森林引导的语义传输用于标签监督的流形对齐

Adrien Aumon, Myriam Lizotte, Guy Wolf, Kevin R. Moon, Jake S. Rhodes

机构 * Department of Mathematics and Statistics, Université de Montréal, Montreal, Canada(蒙特利尔大学数学与统计学系) Mila - Quebec AI Institute, Montreal, Canada(魁北克人工智能研究所) Department of Mathematics and Statistics, Utah State University, Logan, Utah, USA(犹他州立大学数学与统计学系) Department of Statistics, Brigham Young University, Provo, Utah, USA(Brigham Young 大学统计学系)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 FoSTA通过森林引导的语义传输对齐方法,提升多模态数据对齐和标签转移的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏