arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-27 至 2026-02-27 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 9 篇

2602.22405 2026-02-27 cs.LG cs.CV 88%

MolFM-Lite: Multi-Modal Molecular Property Prediction with Conformer Ensemble Attention and Cross-Modal Fusion

MolFM-Lite:基于构象集合注意力和跨模态融合的多模态分子性质预测

Syed Omer Shah, Mohammed Maqsood Ahmed, Danish Mohiuddin Mohammed, Shahnawaz Alam, Mohd Vahaj ur Rahman

机构 * Department of Computer Science and Engineering, University at Buffalo(布法罗大学计算机科学与工程系) Khoury College of Computer Sciences, Northeastern University(东北大学科赫里学院) Department of Computer Science, Muffakham Jah College of Engineering and Technology(穆法卡姆·贾赫工程与技术学院计算机科学系) Department of Computer Science and Artificial Intelligence, Muffakham Jah College of Engineering and Technology(穆法卡姆·贾赫工程与技术学院计算机科学与人工智能系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 MolFM-Lite通过跨模态融合和构象集合注意力机制,提升多模态分子性质预测的准确性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23300 2026-02-27 cs.CL eess.AS 84%

A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations

一种用于对话中多模态情绪识别的专家混合模型

Soumya Dutta, Smruthi Balaji, Sriram Ganapathy

机构 * LEAP Lab, Department of Electrical Engineering(LEAP实验室,电气工程系) Microsoft(微软)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、eess.AS

AI总结 MiSTER-E通过专家混合框架提升对话中多模态情绪识别的准确率,实现跨模态一致性与融合。

Comments Accepted to Elsevier Computer Speech and Language. 30 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15226 2026-02-27 cs.MM cs.CL 84%

Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models

并非所有注意力都是必需的:面向多模态大语言模型的参数和计算高效迁移学习

Qiong Wu, Weihao Ye, Yiyi Zhou, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China(教育部多媒体可信感知与高效计算重点实验室) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CL、cs.MM

AI总结 本文提出高效注意力跳过方法,通过减少冗余注意力计算提升多模态大语言模型的推理效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22917 2026-02-27 cs.CV 83%

Towards Multimodal Domain Generalization with Few Labels

迈向少标签多模态领域泛化的研究

Hongzhao Li, Hao Dong, Hualei Wan, Shupan Li, Mingliang Xu, Muhammad Haris Khan

机构 * Zhengzhou University(郑州大学) ETH Zürich(苏黎世联邦理工学院) MBZUAI(马克斯·普朗克智能系统研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出SSMDG框架,通过三个关键组件实现少标签多模态领域泛化,提升跨模态鲁棒性和领域不变性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22236 2026-02-27 q-bio.GN cs.CV cs.LG 79%

CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction

CrossLLM-Mamba: LLMs多模态状态空间融合用于RNA相互作用预测

Rabeya Tus Sadia, Qiang Ye, Qiang Cheng

机构 * Department of Computer Science, University of Kentucky, Lexington, KY, USA(计算机科学系,肯塔基大学,路易斯维尔,KY,美国) Department of Mathematics, University of Kentucky, Lexington, KY, USA(数学系,肯塔基大学,路易斯维尔,KY,美国) Institute for Biomedical Informatics, University of Kentucky, Lexington, KY, USA(生物医学信息学研究所,肯塔基大学,路易斯维尔,KY,美国)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 CrossLLM-Mamba通过多模态状态空间融合实现RNA相互作用预测,采用双向Mamba编码器和动态序列转换模型,达到高精度性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20723 2026-02-27 cs.AI 79%

Modality-Guided Mixture of Graph Experts with Entropy-Triggered Routing for Multimodal Recommendation

模态引导的图专家混合网络与熵触发路由用于多模态推荐

Ji Dai, Quan Fang, Dengsheng Cai

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tianjin University of Technology(天津理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 MAGNET通过模态引导的图专家混合网络与熵触发路由,提升多模态推荐中融合的可控性、稳定性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22796 2026-02-27 cs.IT math.IT 78%

Multi-modal Data Driven Virtual Base Station Construction for Massive MIMO Beam Alignment

多模态数据驱动的大规模MIMO波束对准虚拟基站构造

Yijie Bian, Wei Guo, Jie Yang, Shenghui Song, Jun Zhang, Shi Jin, Khaled B. Letaief

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

AI总结 本文提出利用多模态数据构建虚拟基站以优化大规模MIMO波束对准,通过几何镜像和稀疏表示提升频谱效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22944 2026-02-27 cs.MM 57%

MViR: Multi-View Visual-Semantic Representation for Fake News Detection

MViR:多视图视觉-语义表示用于虚假新闻检测

Haochen Liang, Xinqi Su, Jun Wang, Chaomeng Chen, Zitong Yu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.MM

AI总结 MViR通过多视图视觉-语义表示方法提升虚假新闻检测的准确性,结合图像多视角特征与文本信息进行融合分析。

Comments Accepted by ICASSP'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22368 2026-02-27 cs.SE cs.AI 57%

EyeLayer: Integrating Human Attention Patterns into LLM-Based Code Summarization

EyeLayer:将人类注意力模式整合到基于LLM的代码摘要中

Jiahao Zhang, Yifan Zhang, Kevin Leach, Yu Huang

机构 * Vanderbilt University(范德比尔特大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 EyeLayer通过整合人类眼动模式提升LLM代码摘要效果,实现13.17%的BLEU-4提升。

Comments Accepted at the 34th IEEE/ACM International Conference on Program Comprehension (ICPC 2026), April 12-13, 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏