arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2605.20449 2026-05-21 cs.LG cs.AI 79%

LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series

LLM预训练塑造了可泛化的流形:跨模态迁移至时间序列的洞察

Alexis Roger, Prateek Humane, Zhenghan Tai, Gwen Legate, Andrei Mircea, Vasilii Feofanov, Irina Rish

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) University of Toronto(多伦多大学) Concordia University(康科迪亚大学) com(42.com)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

AI总结 研究探讨了语言预训练的Transformer能否成为有效的时序预测器,并揭示了跨模态迁移的机制,指出预训练构建了流形,微调则将数值动态投影到任务相关方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18104 2026-05-19 cs.AI cs.CR 79%

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction

多模态大语言模型中的安全几何坍缩与自适应漂移修正

Jiahe Guo, Xiangran Guo, Jiaxuan Chen, Weixiang Zhao, Yanyan Zhao, Yutai Hou, Qianchao Wang, Dandan Tu, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文研究了多模态大语言模型在跨模态安全转移中的不足,提出安全几何坍缩现象,并通过自适应漂移修正方法提升模型安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21941 2026-05-19 cs.LG cs.AI 79%

Robust Multimodal Representation Learning in Healthcare

医疗领域鲁棒多模态表征学习

Xiaoguang Zhu, Linxiao Gong, Lianlong Sun, Yang Liu, Haoyu Wang, Jing Liu

机构 * University of California, Davis(加州大学戴维斯分校) HKUST (GZ)(香港科技大学) University of Rochester(罗切斯特大学) Tongji University(同济大学) Georgia Institute of Technology(佐治亚理工学院) Fudan University(复旦大学) The University of British Columbia(不列颠哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出双流特征去相关框架,通过结构因果分析处理医疗多模态数据中的系统性偏差,提升模型泛化能力,实验验证在MIMIC-IV、eICU和ADNI数据集上的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15460 2026-05-18 cs.IR cs.AI 79%

Differentially Private Motif-Preserving Multi-modal Hashing

差分隐私的动机保持多模态哈希

Zehua Cheng, Wei Dai, Jiahao Sun

机构 * Department of Computer Science\ of Oxford Oxford United Kingdom Department of Computer Science\ of Oxford

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.AI

AI总结 本文提出DMP-MH框架,通过去噪后蒸馏方法在保证隐私的前提下保留多模态数据的结构特征,实验表明其在保持隐私的同时提升了检索性能。

Comments 9 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14406 2026-05-15 cs.LG cs.CV 79%

GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

GeoViSTA:用于多模态环境表示的地理视觉-表格变压器

Yuhao Liu, Sadeer Al-Kindi, Ashok Veeraraghavan, Guha Balakrishnan

机构 * Department of Electrical and Computer Engineering, Rice University(理海大学电气与计算机工程系) Center for Cardiovascular Computational and Precision Health, Department of Cardiology, DeBakey Heart and Vascular Center, Houston Methodist(休斯顿方法主义医疗中心心血管计算与精准健康中心、心内科部门、德贝基心脏和血管中心)

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);分类 cs.CV

AI总结 GeoViSTA通过融合栅格影像与表格数据,构建统一的地理嵌入表示,提升对环境、社会和健康问题的推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14327 2026-05-15 cs.LG cs.AI 79%

AIM-DDI: A Model-Agnostic Multimodal Integration Module for Drug-Drug Interaction Prediction

AIM-DDI: 一种模型无关的多模态整合模块用于药物-药物相互作用预测

Yerin Park, Sangseon Lee

机构 * Department of Artificial Intelligence, Inha University(人工智能系,Inha大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出AIM-DDI模块,通过共享潜在空间整合多模态信息,提升药物-药物相互作用预测的鲁棒性,尤其在未见过的药物情况下表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22853 2026-05-14 cs.CV 79%

Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification

推理时动态模态选择用于不完整多模态分类

Siyi Du, Xinzhe Luo, Declan P. O'Regan, Chen Qin

机构 * Department of Electrical and Electronic Engineering & I-X(电气与电子工程系及I-X)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出DyMo框架,通过动态选择并融合可靠恢复的模态,解决不完整多模态学习中的丢弃或填补困境,实验显示其在多种缺失数据场景下优于现有方法。

Comments 27 pages (including appendix), accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12064 2026-05-13 cs.CV 79%

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

TAR:基于文本语义的跨模态图像配准框架用于光学和SAR图像

Zhuoyu Cai, Dou Quan, Ning Huyan, Pei He, Shuang Wang, Licheng Jiao

机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education of China, School of Artificial Intelligence, Xidian University(中国教育部智能感知与图像理解重点实验室,西安电子科技大学人工智能学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出TAR框架,通过文本语义先验缓解模态差距,提升跨模态特征学习,解决大形变下的光学与SAR图像配准问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11931 2026-05-13 cs.CV 79%

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

学会思考:通过视觉感知的自我改进训练提升多模态推理

Qihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu, Bo Du, Dacheng Tao

机构 * School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence(计算机学院、多媒体软件国家工程研究中心、人工智能研究院) Hubei Key Laboratory of Multimedia(湖北多媒体重点实验室) Network Communication Engineering, Wuhan University, China(网络通信工程、武汉大学,中国) The University of Sydney, Australia(悉尼大学,澳大利亚) Nanyang Technological University, Singapore(南洋理工大学,新加坡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出VISTA框架,通过视觉感知的自我改进训练提升多模态推理能力,解决数据不平衡和语言先验偏差问题,实验显示在多种训练场景下提升性能。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10500 2026-05-13 cs.CV 79%

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning

视觉增强深度缩放用于多模态潜在推理

Yudong Han, Yong Wang, Zaiquan Yang, Zhen Qu, Liyuan Pan, Xiangxiang Chu

机构 * Beijing Institute of Technology(北京理工大学) AMAP, Alibaba Group(阿里集团AMAP) City University of Hong Kong(香港城市大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Yangtze Delta Region Academy of Beijing Institude of Technology, Jiaxing, China(北京理工大学扬子江地区学院,嘉兴,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出视觉回放模块和路由深度缩放,通过增强视觉感知和细化复杂潜在表示,提升多模态潜在推理的效率与性能。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11716 2026-05-13 cs.AI 79%

SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models

SafeSteer: 多模态大语言模型中的解码级防御机制

Xinyi Zeng, Xue Yang, Jingyuan Zhang, Huanqian Yan, Xiang Chen, Kaiwen Wei, Hankun Kang, Yu Tian

机构 * Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Kuaishou Technology(快手科技) School of Computer Science and Technology, Beihang University(北航计算机科学与技术学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Chongqing University(重庆大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出SafeSteer,通过解码阶段的轻量探针和模态语义对齐向量,提升多模态大语言模型的安全性,实验表明其能提升33.40%的安全性而不需微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11015 2026-05-13 cs.CR cs.AI 79%

DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization

DCVD:双通道跨模态融合用于联合漏洞检测与定位

Wenxin Tang, Wenbin Li, Junliang Liu, Jingyu Xiao, Xi Xiao, Mingzhe Liu, Jinlong Yang, Xuan Liu, Yuehe Ma, Wang Luo, Qing Li, Lei Wang, Peng Xiangli

机构 * Tsinghua University(清华大学) Hunan University(湖南大学) Dalian Maritime University(大连海事大学) The Chinese University of Hong Kong(香港中文大学) Shenzhen University(深圳大学) Northwestern Polytechnical University(西北工业大学) Shandong University(山东大学) BNU-HKBU United International College(北京师范大学-香港浸会大学联合国际学院) Sun Yat-sen University(中山大学) Peng Cheng Laboratory(鹏城实验室) Guangzhou Intelligence Communications Technology Co., Ltd.(广州智能通信技术有限公司) The Fifth Electronic Research Institute of MIIT(中华人民共和国信息产业部第五电子研究所)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出DCVD框架,通过双通道融合实现功能级检测与语句级定位的联合优化,有效解决单一信息源和缺乏显式监督的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09614 2026-05-12 cs.CV 79%

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

反射锚点用于长链多模态推理中的传播感知视觉保留

Xuan Gong, Hanbo Huang, Hao Zheng, Yiran Zhang, Wenbin Dai, Weishu Zhao, Shiyu Liang

机构 * Shanghai Jiao Tong University(上海交通大学) Lanzhou University(兰州大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出RAPO方法,通过信息论分析优化视觉传播潜力和局部分支空间,提升长链多模态推理的视觉信息保留效果。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09352 2026-05-12 cs.AI 79%

The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?

维特根斯坦表示假说:语言是否是多模态收敛的吸引子?

Zhaoyang Zhang, Run Shao, Dongyue Wu, Jiajie Teng, Chao Tao, Jingdong Chen, Haifeng Li

机构 * Central South University(中南大学) Huazhong University of Science and Technology(华中科技大学) Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 研究探讨了为何不同模态的神经网络收敛于共享表示,发现语言模态对其他模态有显著方向性吸引,提出维特根斯坦表示假说。

Comments 22 pages, 11 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09151 2026-05-12 cs.CV 79%

MultiMedVision: Multi-Modal Medical Vision Framework

多模态医学视觉框架:MultiMedVision

Frank Li, Bardia Khosravi, Mohammadreza Chavoshi, Young Seok Jeon, Theo Dapamede, Hari Trivedi, Janice Newsome, Judy Gichoya

机构 * Emory University(埃默里大学) Yale University(耶鲁大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出MultiMedVision框架,通过稀疏视觉Transformer实现2D/3D医学影像的统一表征学习,无需模态特定适配器,在共享潜在空间中处理混合模态数据,取得2D和3D任务的竞争力表现。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08764 2026-05-12 cs.LG cs.CV eess.IV 79%

Anchoring the Eigengap: Cross-Modal Spectral Stabilization for Sample-Efficient Representation Learning

锚定特征间隙:跨模态谱稳定化以实现样本高效表征学习

Nikhil J. Dhinagar, Vidhi Chhatbar, Chirag Jagad, Pavithra Senthilkumar, Sophia I. Thomopoulos, Mahir H. Khan, Sook-Lei Liew, the ENIGMA-Stroke Recovery Working Group, Paul M. Thompson

机构 * Imaging Genetics Center, Mark & Mary Stevens Neuroimaging & Informatics Institute, Keck School of Medicine, University of Southern California(影像基因中心,马克与玛丽史蒂文斯神经影像与信息学研究所,凯克医学院,南加州大学) Neuroscience Graduate Program, Mark & Mary Stevens Neuroimaging & Informatics Institute, Chan Division of Occupational Science & Occupational Therapy, Biomedical Engineering, University of Southern California(神经科学研究生项目,马克与玛丽史蒂文斯神经影像与信息学研究所,查恩职业科学与职业治疗 division,生物医学工程,南加州大学)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出跨模态谱稳定化方法,通过抑制噪声主导方向并保持特征间隙,提升样本效率下的表征学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07274 2026-05-11 cs.AI cs.LG 79%

Structured Role-Aware Policy Optimization for Multimodal Reasoning

结构化角色感知策略优化用于多模态推理

Bingqing Jiang, Difan Zou

机构 * School of Computing & Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) School of Computing & Data Science and Institute of Data Science, The University of Hong Kong(计算与数据科学学院和数据科学研究所,香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出结构化角色感知策略优化(SRPO),通过角色感知的token级信用分配提升多模态推理中的证据基础推理能力,无需外部奖励模型。

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06990 2026-05-11 cs.CV cs.LG 79%

TRAJGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations

TRAJGANR: 通过地理对齐的神经表示进行轨迹导向的城市多模态学习

Maria Despoina Siampou, Gengchen Mai, Ni Lao, Jinmeng Rao, Neha Arora, Cyrus Shahabi, Shushman Choudhury

机构 * Google Research, Mountain View, CA(谷歌研究,山景城,加利福尼亚州) Google LLC, Mountain View, CA(谷歌公司,山景城,加利福尼亚州) Dept. of Computer Science, University of Southern California, Los Angeles, CA(计算机科学系,南加州大学,洛杉矶,加利福尼亚州) SEAI Lab, Dept. of Geography and the Environment, The University of Texas at Austin, Austin, TX(SEAI实验室,地理与环境系,德克萨斯大学奥斯汀分校,奥斯汀,德克萨斯州)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TRAJGANR提出一种新的多模态自监督学习框架,通过将连续移动模式与静态位置观察对齐,提升城市理解和移动任务的性能,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19316 2026-05-08 cs.CL 79%

KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

KORE:通过知识导向控制增强大型多模态模型的知识注入

Kailin Jiang, Hongbo Jiang, Ning Jiang, Zhi Gao, Jinhe Bi, Yuchen Ren, Bin Li, Yuntao Du, Lei Liu, Qing Li

机构 * University of Science State Key Laboratory of General Artificial Intelligence, BIGAI Xiamen University Northeast Forestry University Beijing Institute of Technology Ludwig Maximilian University of Munich The University of Sydney C-FAIR\&school of software, Shandong University State Key Lab. for Novel Software Technology, Nanjing University, P.R. China

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 KORE通过知识导向的增强和约束,提升大型多模态模型的知识注入能力,同时保留旧知识。方法利用协方差矩阵和投影初始化,有效减少灾难性遗忘。

Comments ICML 2026, Project Page: https://kore-lmm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05850 2026-05-08 cs.CV 79%

Align3D-AD: Cross-Modal Feature Alignment and Dual-Prompt Learning for Zero-shot 3D Anomaly Detection

Align3D-AD:跨模态特征对齐与双提示学习用于零样本3D异常检测

Letian Bai, Xuanming Cao, Juan Du, Chengyu Tao

机构 * Smart Manufacturing Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)智能制造方向) The Hong Kong University of Science and Technology(香港科技大学) College of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出Align3D-AD框架,通过跨模态特征对齐和双提示学习解决零样本3D异常检测中的领域差距问题,实验表明其在多个数据集上均优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01327 2026-05-08 cs.AI cs.LG 79%

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

基于段落对齐的策略优化用于多模态推理

Lei Gao, Zhuoming Li, Mengxi Jia, Jiakang Yuan, Hongbo Sun, Hao Sun, Xuelong Li

机构 * Fudan University(复旦大学) Southeast University(东南大学) China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.(中国电信人工智能技术(北京)有限公司) Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出SAPO方法,通过将推理步骤而非token或完整序列作为策略更新的基本单元,提升多模态推理任务的准确性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01061 2026-05-05 cs.MM 79%

PRISM: Exposing and Resolving Spurious Isolation in Federated Multimodal Continual Learning

PRISM:揭示并解决联邦多模态持续学习中的虚假隔离

Beining Wu, Zihao Ding, Jun Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出PRISM方法,通过维护专家梯度子空间基底和重新解释MoE路由,解决联邦多模态持续学习中路由隔离失效、遗忘积累和梯度冲突问题,实验表明其在多个数据集上优于现有基线。

Comments submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00973 2026-05-05 cs.LG cs.AI eess.SP 79%

Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning

生理感知的跨模态掩码重建用于生物信号表示学习

Hao Zhou, Simon A. Lee, Cyrus Tanade, Keum San Chun, Juhyeon Lee, Migyeong Gwak, Megha Thukral, Justin Sung, Eugene Hwang, Mehrab Bin Morshed, Li Zhu, Viswam Nathan, Md Mahbubur Rahman, Subramaniam Venkatraman, Sharanya Arcot Desai

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Samsung Research America(三星美国研究院)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出xMAE框架,通过跨模态掩码重建捕捉生物信号的时序关系,提升表示学习效果,在15/19下游任务中优于基线模型。

Comments Proceedings of the 43rd International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10554 2026-05-05 cs.CV 79%

Grounding Everything in Tokens for Multimodal Large Language Models

基于令牌的多模态大语言模型的 grounding

Xiangxuan Ren, Zhongdao Wang, Liping Hou, Pin Tang, Guoqing Wang, Chao Ma

机构 * MoE Key Lab of Artificial Intelligence(人工智能混合专家模型关键实验室) AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能研究院) Central Research Institute, Huawei(华为中央研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出GETok方法,通过整合可学习的令牌词汇提升多模态大语言模型在2D空间中的物体定位能力,实验显示其在多种指引用例任务中表现更优。

Comments 19 pages, 16 figures, 12 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10320 2026-05-05 cs.CV eess.IV 79%

Boosting Multimodal Remote Sensing Image Classification with Transformer-based Heterogeneously Salient Graph Representation

通过基于变换器的异质显著图表示提升多模态遥感图像分类

Jiaqi Yang, Bo Du, Rong Liu, Zhu Mao, Liangpei Zhang

机构 * Department of Forest Sciences, University of Helsinki(赫尔辛基大学森林科学系) School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, Wuhan, China(武汉大学计算机学院、国家多媒体软件工程研究中心、人工智能研究院、湖北省多媒体与网络通信工程重点实验室,武汉,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于变换器的异质显著图表示方法,解决多模态遥感图像分类中特征表示不充分、长距离依赖建模复杂及过拟合问题,通过异质图编码器、多卷积调节器和均值前向策略提升分类性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04465 2026-05-04 cs.AI 79%

The Topology of Multimodal Fusion: Why Current Architectures Fail at Creative Cognition

多模态融合的拓扑学:为何当前架构在创造性认知上失败

Xiujiang Tan

机构 * Guangzhou Academy of Fine Arts(广州美术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文指出当前多模态AI架构存在拓扑限制而非参数限制,通过哲学、认知科学和数学三个支柱提出多模态融合的拓扑学框架,提出UOO实现、ANALOGY-MM基准和META-TOP三阶基准,通过实验路线验证拓扑同构性。

Comments Expanded 11 technical improvements; 5 reference corrections; Appendix B pseudocode added. ~43 pages, 5 figures. Chinese philosophical terms romanized. Companion monograph available separately

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26067 2026-04-30 cs.CV 79%

RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

RADIO-ViPE:面向动态环境的在线紧密耦合多模态融合用于开放词汇语义SLAM

Zaid Nasser, Mikhail Iumanov, Tianhao Li, Maxim Popov, Jaafar Mahmoud, Sergey Kolyubin

机构 * Biomechatronics and Energy-Efficient Robotics (BE2R) Lab, ITMO University(生物机电学与高效能机器人实验室,ITMO大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 RADIO-ViPE通过在线多模态融合实现动态环境中开放词汇语义SLAM,无需校准RGB-D输入,结合视觉与语言嵌入提升地图一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12052 2026-04-29 cs.CV 79%

Task-Driven Prompt Learning: A Joint Framework for Multi-modal Cloud Removal and Segmentation

任务驱动的提示学习:多模态云去除与分割的联合框架

Zaiyan Zhang, Jie Li, Shaowei Shi, Qiangqiang Yuan

机构 * Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出TDP-CR框架,通过任务驱动的多模态方法联合实现云去除与土地覆盖分割,解决云遮挡导致的语义不一致问题,提升分析级数据质量。

Comments Accepted by IGARSS 2026 Conference (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14245 2026-04-29 cs.LG cond-mat.mtrl-sci cs.AI cs.CE q-bio.BM 79%

Curriculum-guided multimodal representation learning enables generalizable prediction of nanomaterial-protein interactions

基于课程引导的多模态表示学习可实现纳米材料-蛋白质相互作用的通用预测

Hengjie Yu, Kenneth A. Dawson, Haiyun Yang, Shuya Liu, Yan Yan, Yaochu Jin

机构 * School of Engineering, Westlake University(西湖大学工程学院) Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖研究所在先进科技研究院) Centre for BioNano Interactions, School of Chemistry, University College Dublin(都柏林大学学院化学系生物纳米相互作用中心) School of Biomolecular and Biomedical Science, UCD Conway Institute of Biomolecular and Biomedical Research(都柏林大学学院生物分子与生物医学科学系,康沃利斯生物分子与生物医学研究学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出CuMMI模型,通过多阶段课程学习和大规模数据集,实现纳米材料-蛋白质相互作用的通用预测,验证了其在不同数据集上的鲁棒性和可迁移性。

Comments 36 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23753 2026-04-28 cs.AI cs.HC cs.LG 79%

Modeling Induced Pleasure through Cognitive Appraisal Prediction via Multimodal Fusion

通过多模态融合进行诱导愉悦的认知评估预测建模

Nastaran Dab, Raziyeh Zall, Mohammadreza Kangavari

机构 * Iran University of Science and Technology(伊朗科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种多模态融合模型,通过认知评估变量预测视频诱导的愉悦,解决标签噪声、语义鸿沟、数据稀缺和解释性不足等问题,实验验证了模型在检测视频诱导愉悦方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏