arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2603.17753 2026-03-19 cs.CV 79%

PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation

PC-CrossDiff:点-簇双级跨模态微分注意力用于统一的3D指称与分割

Wenbin Tan, Jiawen Lin, Fangyong Wang, Yuan Xie, Yong Xie, Yachao Zhang, Yanyun Qu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出PC-CrossDiff框架,通过双级跨模态微分注意力解决复杂多物体场景中指称理解和分割的挑战,提升3D视觉定位的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17347 2026-03-19 cs.MM 79%

Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning

超越强制模态平衡:多模态学习中的内在信息预算

Zechang Xiong, Da Li, Kexin Tang, Pengyuan Li, Wenkang Kong, Yulan Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出IIBalance框架,通过内在信息预算对齐模态贡献,解决多模态学习中的模态不平衡问题,实验表明其优于现有方法。

Comments 6 pages, 4 figures, paper accepted by ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16259 2026-03-18 cs.MM 79%

Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction

双曲多模态生成表示学习用于广义零样本多模态信息提取

Baohang Zhou, Kehui Song, Rize Jin, Yu Zhao, Xuhui Sui, Xinying Qian, Xingyue Guo, Ying Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出双曲多模态生成表示学习框架HMGRL,解决零样本多模态信息提取中见与不见类别共存的问题,通过双曲空间建模多级语义关联,提升模型泛化能力。

Comments Accepted by WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16143 2026-03-18 eess.SP cs.AI 79%

Structure-Aware Multimodal LLM Framework for Trustworthy Near-Field Beam Prediction

面向可信近场波束预测的结构感知多模态大语言模型框架

Mengyuan Li, Qianfan Lu, Jiachen Tian, Hongjun Hu, Yu Han, Xiao Li, Chao-kai Wen, Shi Jin

机构 * School of Information Science and Engineering, Southeast University(信息科学与工程学院,东南大学) Institute of Communications Engineering, National Sun Yat-sen University(通讯工程学院,国立中山大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出基于大语言模型的多模态框架,融合历史GPS数据、RGB图像、LiDAR数据及任务特定文本提示,以提升复杂低空环境中的近场波束对齐能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15100 2026-03-17 cs.CV 79%

Learning from Limited and Incomplete Data: A Multimodal Framework for Predicting Pathological Response in NSCLC

从有限和不完整数据中学习:一种多模态框架用于预测非小细胞肺癌的病理反应

Alice Natalina Caragliano, Giulia Farina, Fatih Aksu, Camillo Maria Caruso, Claudia Tacconi, Carlo Greco, Lorenzo Nibid, Edy Ippolito, Michele Fiore, Giuseppe Perrone, Sara Ramella, Paolo Soda, Valerio Guarrasi

机构 * Department of Medicine and Surgery(医学与外科系) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering(诊断与介入系,放射物理,生物医学工程) Umeå University, Umeå, Sweden(乌梅大学,瑞典)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出一种多模态深度学习框架,整合基础模型的CT特征提取与缺失感知架构,以在有限数据和不完整临床资料下准确预测NSCLC的病理反应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13298 2026-03-17 cs.LG cs.AI 79%

FusionCast: Enhancing Precipitation Nowcasting with Asymmetric Cross-Modal Fusion and Future Radar Priors

FusionCast: 通过不对称跨模态融合和未来雷达先验增强降水现在预报

Henan Wang, Shengwu Xiong, Yifang Zhang, Wenjie Yin, Chen Zhou, Yuqiang Zhang, Pengfei Duan

机构 * School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院) School of Earth and Space Science and Technology, Wuhan University(武汉大学地球和空间科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出FusionCast框架,结合历史雷达QPE、PWV数据和预测雷达QPE,通过不对称融合和未来先验提升降水现在预报精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13291 2026-03-17 cs.LG cs.AI 79%

FedUAF: Uncertainty-Aware Fusion with Reliability-Guided Aggregation for Multimodal Federated Sentiment Analysis

FedUAF: 基于不确定性融合与可靠性引导聚合的多模态联邦情感分析

Xianxun Zhu, Zezhong Sun, Imad Rida, Erik Cambria, Junqi Su, Rui Wang, Hui Chen

机构 * Shanghai University(上海大学) North China Electric Power University(华北电力大学) Université de Technologie de Compiègne(法国图卢兹国立理工学院) Nanyang Technological University(南洋理工大学) City University of Hong Kong(香港城市大学) Macquarie University(麦考瑞大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出FedUAF框架,通过不确定性融合和可靠性引导聚合解决联邦学习中多模态数据缺失、分布异质和客户端更新不可靠的问题,实验证明其在CMU-MOSI和CMU-MOSEI数据集上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12762 2026-03-16 cs.CV cs.LG 79%

TerraFlow: Multimodal, Multitemporal Representation Learning for Earth Observation

TerraFlow:用于地球观测的多模态、多时间序列表示学习

Nazar Puriy, Johannes Jakubik, Benedikt Blumenstiel, Konrad Schindler

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TerraFlow提出了一种新的多模态、多时间序列学习方法,适用于地球观测,能够处理变长输入,优于现有模型,在GEO-Bench-2基准测试中表现优异,并在自然灾害风险地图预测中取得进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12166 2026-03-13 cs.CV 79%

LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning

LatentGeo: 在潜在空间中学习可学习的辅助构造以进行多模态几何推理

Haiying Xu, Zihan Wang, Song Dai, Zhengxuan Zhang, Kairan Dou, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nankai University(南开大学) Communication University of China(中国传媒大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 LatentGeo通过学习潜在空间中的连续视觉表示,解决多模态几何推理中辅助构造的表示问题,采用三阶段课程和强化学习方法提升几何推理任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02686 2026-03-13 cs.LG cs.AI 79%

A Systematic Review of Intermediate Fusion in Multimodal Deep Learning for Biomedical Applications

多模态深度学习在生物医学应用中的中间融合系统综述

Valerio Guarrasi, Fatih Aksu, Camillo Maria Caruso, Francesco Di Feola, Aurora Rofena, Filippo Ruffini, Paolo Soda

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文系统综述了多模态深度学习在生物医学应用中的中间融合方法,分析了现有技术、挑战及未来方向,并提出结构化符号以促进方法的广泛应用。

Journal ref Image and Vision Computing 158 (2025) 105509

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09706 2026-03-11 cs.AI 79%

OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences

OOD-MMSafe: 推进多模态大语言模型安全性从有害意图到隐藏后果

Ming Wen, Kun Yang, Jingyu Zhang, Yuxuan Liu, shiwen cui, Shouling Ji, Xingjun Ma

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.AI

AI总结 OOD-MMSafe通过CASPO框架提升多模态大语言模型对隐藏后果的识别能力,显著降低风险识别失败率。

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09258 2026-03-11 cs.CV 79%

Multimodal Graph Representation Learning with Dynamic Information Pathways

多模态图表示学习中的动态信息路径

Xiaobin Hong, Mingkai Lin, Xiaoli Wang, Chaoqun Wang, Wenzhong Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出DiP框架,通过动态信息路径实现多模态图表示学习,提升跨模态消息传播的适应性和表达性。

Comments 12 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13015 2026-03-11 cs.CV 79%

Multimodal Classification via Total Correlation Maximization

通过总相关性最大化实现多模态分类

Feng Yu, Xiangyu Wu, Yang Yang, Jianfeng Lu

机构 * Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出TCMax方法,通过最大化多模态特征与标签间的总相关性,缓解模态竞争并提升多模态分类性能。

Comments Accepted for publication at ICLR 2026; 19 pages; 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21064 2026-03-11 cs.CV 79%

Multimodal Skeleton-Based Action Representation Learning via Decomposition and Composition

多模态骨骼基动作表示学习通过分解与组合

Hongsong Wang, Heng Fei, Bingxuan Dai, Jie Gui

机构 * School of Computer Science and Engineering, Southeast University, Nanjing 210096, China(东南大学计算机科学与工程学院,南京210096,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用关键实验室(东南大学),中华人民共和国教育部,中国) School of Cyber Science and Engineering, Southeast University, Nanjing 210096, China(东南大学网络安全科学与工程学院,南京210096,中国) Engineering Research Center of Blockchain Application, Supervision And Management (Southeast University), Ministry of Education, China(区块链应用、监督与管理工程研究中心(东南大学),中华人民共和国教育部,中国) Purple Mountain Laboratories, Nanjing 210000, China(紫金山实验室,南京210000,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出一种自监督的多模态骨骼基动作表示学习框架,通过分解与组合策略平衡效率与效果,提升动作识别性能。

Comments Accepted by Machine Intelligence Research (Journal Impact Factor 8.7, 2024)

Journal ref Machine Intelligence Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07874 2026-03-10 cs.CV cs.LG 79%

Toward Unified Multimodal Representation Learning for Autonomous Driving

迈向自动驾驶的统一多模态表示学习

Ximeng Tao, Dimitar Filev, Gaurav Pandey

机构 * J. Mike Walker ’66 Department of Mechanical Engineering, Texas A&M University, College Station, TX 77843, USA(德克萨斯大学机械工程系,德克萨斯农工大学,学院站,德克萨斯,77843,美国) The Department of Engineering Technology and Industrial Distribution Texas A&M University, College Station, TX 77843, USA(工程技术与工业分布系,德克萨斯农工大学,学院站,德克萨斯,77843,美国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CTP框架,通过统一多模态张量对齐提升自动驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10328 2026-03-09 cs.CV 79%

Fuse4Seg: Image Fusion for Multi-Modal Medical Segmentation via Bi-level Optimization

Fuse4Seg: 多模态医学分割的图像融合 via 两级优化

Yuchen Guo, Junli Gong, Hongmin Cai, Yiu-ming Cheung, Weifeng Su

机构 * Northwestern University(西北大学) Northeastern University(东北大学) South China University of Technology(华南理工大学) Hong Kong Baptist University(香港 Baptist大学) Beijing Normal - Hong Kong Baptist University(北京师范大学-香港 Baptist大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 Fuse4Seg通过两级优化实现多模态医学图像融合,解决视觉与语义间的差距问题,提升分割任务的准确性和临床可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04887 2026-03-06 cs.CV 79%

Federated Modality-specific Encoders and Partially Personalized Fusion Decoder for Multimodal Brain Tumor Segmentation

联邦模态特定编码器和部分个性化融合解码器用于多模态脑肿瘤分割

Hong Liu, Dong Wei, Qian Dai, Xian Wu, Yefeng Zheng, Liansheng Wang

机构 * National Institute for Data Science in Health and Medicine(国家医学数据科学研究院) Department of Computer Science at School of Informatics(信息学院计算机科学系) Jarvis Research Center(Jarvis研究中心) Medical Artificial Intelligence Lab(医学人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出FedMEPD框架,通过联邦模态特定编码器和部分个性化融合解码器,解决多模态医学图像分析中的模态间异质性和个性化需求问题。

Comments Medical Image Analysis 2025. arXiv admin note: substantial text overlap with arXiv:2403.11803

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04562 2026-03-06 cs.CV cs.LG 79%

Fusion and Grouping Strategies in Deep Learning for Local Climate Zone Classification of Multimodal Remote Sensing Data

深度学习中多模态遥感数据局部气候区分类的融合与分组策略

Ancymol Thomas, Jaya Sreevalsan-Nair

机构 * Graphics-Visualization-Computing Lab, International Institute of Information Technology Bangalore, Karnataka 560100, India(图形可视化计算实验室,国际信息学院班加罗尔,卡纳塔克邦560100,印度)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出了一种基于深度学习的多模态遥感数据局部气候区分类方法,通过融合与分组策略提升分类准确率,最终达到76.6%的整体准确率。

Comments 25 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02629 2026-03-04 cs.CV 79%

Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective

迈向增量统一多模态异常检测:从信息瓶颈视角增强多模态去噪

Kaifang Long, Lianbo Ma, Jiaqi Liu, Liming Liu, Guoyang Xie

机构 * Software College, Northeastern University, China(东北大学软件学院) CATL, China(宁德时代)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出IB-IUMAD框架,通过Mamba解码器和信息瓶颈融合模块解决多模态异常检测中的灾难性遗忘问题,提升模型对新兴对象的适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02162 2026-03-03 cs.CV 79%

Bridging the gap between Performance and Interpretability: An Explainable Disentangled Multimodal Framework for Cancer Survival Prediction

弥合性能与可解释性之间的鸿沟:一种可解释的解耦多模态框架用于癌症生存预测

Aniek Eijpe, Soufyan Lakbir, Melis Erdal Cesur, Sara P. Oliveira, Angelos Chatzimparmpas, Sanne Abeln, Wilson Silva

机构 * AI Technology for Life(人工智能技术与生命科学) Department of Information and Computing Sciences(信息与计算科学系) Department of Biology(生物学系) Utrecht University(乌得勒支大学) Department of Metabolic Diseases(代谢疾病部门) Wilhelmina Children’s Hospital(维廉明娜儿童医院) University Medical Center Utrecht(乌得勒支大学医学中心) Regenerative Medicine Center Utrecht(乌得勒支再生医学中心) Computational Pathology(计算病理学) Department of Pathology(病理学系) The Netherlands Cancer Institute(荷兰癌症研究所) Visualization and Graphics(可视化与图形学) The Netherlands Cancer Institute, Amsterdam(荷兰癌症研究所,阿姆斯特丹)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 DIMAFx通过解耦多模态表示提升癌症生存预测的性能与可解释性,揭示了多模态交互和生物学信息。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01758 2026-03-03 cs.CV 79%

Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining

通过语言枢轴预训练统一异构多模态遥感检测

Yuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng, Xiang Li, Jian Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 BabelRS通过语言枢轴预训练框架统一异构多模态遥感检测,解耦模态对齐与任务学习,提升训练稳定性与检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09285 2026-03-03 cs.CV 79%

Spotlight on Token Perception for Multimodal Reinforcement Learning

多模态强化学习中的token感知聚焦

Siyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo, Zefeng He, Daizong Liu, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University(南京大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出VPPO算法,通过token感知优化提升多模态强化学习的视觉推理能力。

Comments Accepted by ICLR 2026, project page: https://github.com/huaixuheqing/VPPO-RL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03214 2026-03-03 cs.CV 79%

RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion

RTGMFF:基于ROI驱动文本生成和多模态特征融合的增强型fMRI脑部疾病诊断

Junhao Jia, Yifei Sun, Yunyou Liu, Cheng Yang, Changmiao Wang, Feiwei Qin, Yong Peng, Wenwen Min

机构 * Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University(浙江大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) Yunnan University(云南大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 RTGMFF通过结合ROI驱动文本生成和多模态特征融合,提升fMRI在脑部疾病诊断中的准确性。

Comments The paper has been accepted by BIBM 2025

Journal ref 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE, 2025, pp. 2301-2308

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22283 2026-03-03 cs.CV 79%

Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment

重新审视在跨模态不匹配下的LVLMs视觉标记减少

Rui Xu, Yunke Wang, Yong Luo, Bo Du

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出VisionDrop方法,通过视觉-only修剪框架减少LVLMs中的视觉标记,无需额外训练,提升推理效率并保持性能。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00157 2026-03-03 cs.CV 79%

FujiView: Multimodal Late-Fusion for Predicting Scenic Visibility

FujiView: 多模态晚期融合用于预测风景可见性

Bryceton Bible, Shah Md Nehal Hasnaeen, Hairong Qi

机构 * University of Tennessee, Knoxville(田纳西大学,科文克顿)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 FujiView通过融合摄像头图像与气象数据,实现风景可见性的多模态预测,展示了在短期和长期预测中的不同方法效果。

Comments 9 pages (including references), 8 figures, 2 tables. Accepted to the IEEE/CVF WACV 2026 proceedings. Introduces a large human-labeled Mount Fuji visibility dataset; public release forthcoming

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24041 2026-03-02 cs.CV 79%

Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation

仔细观察:多模态大语言模型中的自适应视觉增强以缓解幻觉

Xingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang, Zhicai Wang, Beier Zhu, Hanwang Zhang

机构 * MoE Key Lab of BIPC, University of Science and Technology of China(信息与电子技术联合实验室,中国科学技术大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出自适应视觉增强框架AIR,通过减少冗余标记和选择性整合补丁来缓解多模态大语言模型中的幻觉问题。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04576 2026-03-02 cs.CV 79%

TARDis: Time Attenuated Representation Disentanglement for Incomplete Multi-Modal Tumor Segmentation and Classification

TARDis: 用于不完整多模态肿瘤分割与分类的时间衰减表示解耦

Zishuo Wan, Qinqin Kang, Na Li, Yi Huang, Qianru Zhang, Le Lu, Yun Bian, Dawei Ding, Ke Yan

机构 * School of Automation and Electrical Engineering, University of Science and Technology Beijing(北京科技大学自动化与电气工程学院) Alibaba Group DAMO Academy(阿里巴巴集团DAMO学院) Hupan Lab(湖畔实验室) Departments of Radiology, Changhai Hospital(上海长海医院放射科)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 TARDis通过时间衰减表示解耦框架,解决多模态肿瘤分割与分类中缺失模态问题,提升诊断精度并降低辐射暴露。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17966 2026-03-02 cs.IR cs.CV 79%

LLM-Enhanced Multimodal Fusion for Cross-Domain Sequential Recommendation

基于大语言模型的跨域序列推荐多模态融合

Wangyu Wu, Zhenhong Chen, Wenqiao Zhang, Xianglin Qiu, Siqi Song, Xiaowei Huang, Fei Ma, Jimin Xiao

机构 * Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Liverpool(利物浦大学) Microsoft(微软公司) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出LLM-EMF方法,通过融合视觉和文本数据提升跨域序列推荐性能,利用CLIP模型生成多模态嵌入并引入多重注意力机制,实验证明其在多领域推荐中的有效性。

Comments arXiv admin note: substantial text overlap with arXiv:2504.15085

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01728 2026-03-02 cs.CV 79%

Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion

Shuffle Mamba:基于随机洗牌的态空间模型用于多模态图像融合

Ke Cao, Xuanhua He, Tao Hu, Chengjun Xie, Man Zhou, Jie Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Intelligent Machines(智能机器研究所) Hefei Institutes of Physical Science, Chinese Academy of Sciences(中国科学院合肥物质科学研究院) Intelligent Agriculture Engineering Laboratory of Anhui Province, Institute of Intelligent Machines(安徽省智能农业工程实验室,智能机器研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 Shuffle Mamba通过引入随机洗牌策略和逆洗牌,解决多模态图像融合中固定扫描策略带来的偏见问题,提升融合质量。

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22236 2026-02-27 q-bio.GN cs.CV cs.LG 79%

CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction

CrossLLM-Mamba: LLMs多模态状态空间融合用于RNA相互作用预测

Rabeya Tus Sadia, Qiang Ye, Qiang Cheng

机构 * Department of Computer Science, University of Kentucky, Lexington, KY, USA(计算机科学系,肯塔基大学,路易斯维尔,KY,美国) Department of Mathematics, University of Kentucky, Lexington, KY, USA(数学系,肯塔基大学,路易斯维尔,KY,美国) Institute for Biomedical Informatics, University of Kentucky, Lexington, KY, USA(生物医学信息学研究所,肯塔基大学,路易斯维尔,KY,美国)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 CrossLLM-Mamba通过多模态状态空间融合实现RNA相互作用预测,采用双向Mamba编码器和动态序列转换模型,达到高精度性能。

详情

展开后加载摘要…

URL PDF HTML 收藏