arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-01 至 2025-12-01 共收录 111 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 18 篇

2511.23287 2025-12-01 cs.LG cs.CL 83%

Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla

基于Transformer的三融合框架用于低资源孟加拉语多模态作者意图分类

Ariful Islam, Tanvir Mahmud, Md Rifat Hossen

机构 * Department of Computer Science(计算机科学系) Engineering Chittagong University of Engineering(工程学院恰尔达格工程大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出基于Transformer的三融合框架BangACMM,通过结合文本和视觉数据,在低资源孟加拉语社交媒体中实现作者意图分类,达到84.11%的宏F1得分,提升8.4个百分点。

Comments Accepted at the 28th International Conference on Computer and Information Technology (ICCIT 2025). To be published in IEEE proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22897 2025-12-01 cs.CV 83%

From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts

从点到云:学习多模态提示的稳健语义分布

Weiran Li, Yeqiang Liu, Yijie Wei, Mina Han, Xin Liu, Zhenbo Li

机构 * China Agricultural University(中国农业大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出P2C框架,通过动态去噪机制学习语义云分布,提升多模态提示学习的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22103 2025-12-01 cs.CV 83%

MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding

MoE3D:专家混合方法与多模态3D理解

Yu Li, Yuenan Hou, Yingmei Wei, Xinge Zhu, Yuexin Ma, Wenqi Shao, Yanming Guo

机构 * National University of Defense Technology(国防科技大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) ShanghaiTech University(上海科技大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MoE3D通过整合专家混合方法,提升多模态3D理解的性能,尤其在Multi3DRefer任务中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01528 2025-12-01 cs.CL cs.AI q-bio.BM 81%

Leveraging Biomolecule and Natural Language through Multi-Modal Learning: A Survey

利用生物分子和自然语言通过多模态学习:一篇综述

Qizhi Pei, Zhimeng Zhou, Kaiyuan Gao, Jinhua Zhu, Yue Wang, Zun Wang, Tao Qin, Lijun Wu, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) Huazhong University of Science and Technology(华中科技大学) University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Shanghai AI Laboratory(上海人工智能实验室) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文综述了生物分子与自然语言多模态学习的最新进展,探讨了技术表示、多模态整合方法、应用实例及未来研究方向。

Comments 2025.11.28 Updated Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22961 2025-12-01 cs.CV 79%

HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model

HMR3D:用于大视觉-语言模型的层次多模态表示以实现3D场景理解

Chen Li, Eric Peh, Basura Fernando

机构 * Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡) Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore(前沿人工智能研究中心,科技研究局,新加坡) College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 HMR3D通过层次化多模态表示,结合多视图图像和文本描述,提升3D场景理解的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21937 2025-12-01 cs.CV 79%

Interpretable Multimodal Cancer Prototyping with Whole Slide Images and Incompletely Paired Genomics

可解释的多模态癌症原型生成:结合整张滑片图像与不完全配对的基因组学

Yupei Zhang, Yating Huang, Wanming Hu, Lequan Yu, Hujun Yin, Chao Li

机构 * Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) Department of Electrical & Electronic Engineering, The University of Manchester, UK(曼彻斯特大学电气与电子工程系) Department of Pathology, State Key Laboratory of Oncology in South China, Guangdong Provincial Clinical Research Center for Cancer, Sun Yat-sen University Cancer Center, China(南方医科大学肿瘤学国家重点实验室、广东省癌症临床研究中心、中山大学肿瘤中心病理学部) Department of Statistics and Actuarial Science, The University of Hong Kong, Hong Kong SAR, China(香港大学统计与精算科学系) Department of Clinical Neurosciences and Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学临床神经科学系和应用数学与理论物理系;邓迪大学科学与工程学院和医学学院) School of Science and Engineering and School of Medicine, University of Dundee, UK

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种可解释的多模态原型生成框架,通过整合整张滑片图像和不完整的基因组学数据,提升精准肿瘤学中的多模态整合效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12099 2025-12-01 cs.CV 79%

TinyRS-R1: Compact Multimodal Language Model for Remote Sensing

TinyRS-R1:用于遥感的紧凑多模态语言模型

Aybora Koksal, A. Aydin Alatan

机构 * Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中欧技术大学(METU)电子与电气工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TinyRS-R1是一种专为遥感设计的紧凑多模态语言模型,通过四阶段训练实现高效性能,兼具推理增强与低资源消耗。

Comments Accepted to IEEE Geoscience and Remote Sensing Letters (GRSL). Code, models, and the captions for datasets are available at https://github.com/aybora/TinyRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01858 2025-12-01 quant-ph cs.CR cs.DC cs.ET cs.LG 78%

MQFL-FHE: Multimodal Quantum Federated Learning Framework with Fully Homomorphic Encryption

MQFL-FHE: 多模态量子联邦学习框架与全同态加密

Siddhant Dutta, Nouhaila Innan, Sadok Ben Yahia, Muhammad Shafique, David Esteban Bernal Neira

机构 * SVKM's Dwarkadas J. Sanghvi College of Engineering(SVKM的达沃拉德·J·桑格维工程学院) eBRAIN Lab, Division of Engineering, New York University Abu Dhabi (NYUAD)(eBRAIN实验室,工程系,纽约大学阿布扎比分校(NYUAD)) Center for Quantum and Topological Systems (CQTS), NYUAD Research Institute(量子与拓扑系统中心(CQTS),NYUAD研究机构) The Maersk Mc-Kinney Moller Institute, University of Southern Denmark(马士克·麦克金尼·莫勒研究所,南丹麦大学) Corvinus Institute for Advanced Studies (CIAS), Budapest, Hungary(科维努斯高级研究学院(CIAS),布达佩斯,匈牙利) Davidson School of Chemical Engineering, Purdue University(戴维森化学工程学院,普渡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出一种结合量子计算和全同态加密的多模态联邦学习框架,旨在提升隐私保护下的模型性能与泛化能力。

Comments 10 pages, 6 figures, 6 Tables. Accepted at IJCNN 2025

Journal ref 2025 International Joint Conference on Neural Networks (IJCNN), Rome, Italy, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00964 2025-12-01 eess.SP 78%

Multi-Modal Fusion-Based Multi-Task Semantic Communication System

基于多模态融合的多任务语义通信系统

Zengle Zhu, Rongqing Zhang, Xiang Cheng, Liuqing Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

AI总结 本文提出基于多模态融合的多任务语义通信系统,利用BERT融合模块提升多模态信息处理效率,实验表明其在性能和通信开销上优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21889 2025-12-01 cs.LG 78%

Exploring Fusion Strategies for Multimodal Vision-Language Systems

探索多模态视觉-语言系统的融合策略

Regan Willis, Jason Bakos

机构 * University of South Carolina(南卡罗来纳大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文探讨了多模态视觉-语言系统中不同数据融合策略的准确率与延迟权衡,通过三种模型架构验证了早期融合在降低延迟方面的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04096 2025-12-01 cs.CE 78%

Cross-Modal Alignment between Visual Stimuli and Neural Responses in the Visual Cortex

视觉皮层中视觉刺激与神经反应的跨模态对齐

Xing Gao, Dazhong Rong, Qinming He

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

AI总结 本文提出视觉-神经对齐方法VNA,通过判别性编码和解码任务在视觉皮层中实现更精确的视觉刺激与神经反应映射。

Comments This paper has been accepted by 2025 International Conference on Brain-Computer Interface (ICBCI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21395 2025-12-01 cs.CV cs.AI 62%

Monet: Reasoning in Latent Visual Space Beyond Images and Language

Monet: 在图像和语言之外的潜在视觉空间推理

Qixun Wang, Yang Shi, Yifei Wang, Yuanxing Zhang, Pengfei Wan, Kun Gai, Xianghua Ying, Yisen Wang

机构 * Peking University(北京大学) Kling Team(Kling团队) Amazon AGI SF Lab(Amazon AGI SF实验室) MIT(麻省理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Monet通过在潜在视觉空间中直接推理,提升多模态大语言模型的视觉推理能力,解决了潜在-视觉对齐和嵌入监督不足的问题,展示了在现实世界和抽象视觉任务中的优越表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22172 2025-12-01 cs.CV 57%

Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning

引导内在之眼:一种用于层次化和灵活的视觉基础推理的框架

Zhaoyang Wei, Wenchao Ding, Yanchao Hao, Xi Chen

机构 * Basic Algorithm Center, PCG, Tencent(腾讯基本算法中心、PCG、腾讯)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 GRiP通过认知增强的强化学习框架,提升视觉基础推理的鲁棒性和灵活性,实现复杂场景下的高性能表现。

Comments 9pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22107 2025-12-01 cs.CV 57%

HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction

HyperST:面向空间转录组预测的分层双曲学习

Chen Zhang, Yilu An, Ying Chen, Hao Li, Xitong Ling, Lihao Liu, Junjun He, Yuxiang Lin, Zihui Wang, Rongshan Yu

机构 * Xiamen University(厦门大学) Shanghai AI Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 HyperST通过在双曲空间中建模数据的层次结构,实现空间转录组预测的多级图像-基因表示学习,提升跨模态预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21741 2025-12-01 cs.CL cs.LG stat.ML 57%

A Multiscale Geometric Method for Capturing Relational Topic Alignment

一种多尺度几何方法用于捕捉关系主题对齐

Conrad D. Hougen, Karl T. Pazdernik, Alfred O. Hero

机构 * University of Michigan(密歇根大学) Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出一种多尺度几何方法,结合多模态文本和合著者网络数据,以捕捉关系主题对齐并识别罕见话题结构。

Comments 5 pages, 3 figures, 2025 IEEE International Workshop on Computational Advances in Multi-Sensor Adaptive Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22859 2025-12-01 eess.IV cs.CR 50%

TokCom-UEP: Semantic Importance-Matched Unequal Error Protection for Resilient Image Transmission

TokCom-UEP:语义重要性匹配的非等效保护用于稳健图像传输

Kaizheng Zhang, Zuolin Jin, Zhihang Cheng, Ming Zeng, Li Qiao, Zesong Fei

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 TokCom-UEP通过语义重要性匹配的非等效保护提升图像传输的稳健性与效率

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13348 2025-12-01 cond-mat.soft 50%

Power Law Rheology of Folded Protein Hydrogels

折叠蛋白质水凝胶的幂律流变学

Anders Aufderhorst-Roberts, Sophie Cussons, David J. Brockwell, Lorna Dougan

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 研究显示,折叠蛋白质水凝胶在mesoscopic尺度上表现出独特的幂律粘弹性特性,揭示了其粘弹性质的多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 4 篇

2502.19674 2025-12-01 cs.CV 79%

Reliable Multimodal Learning Via Multi-Level Adaptive DeConfusion

通过多级自适应去混淆实现可靠的多模态学习

Tong Zhang, Shu Shen, C. L. Philip Chen

机构 * Guangdong Provincial Key Laboratory of Computational AI Models and Cognitive Intelligence(广东省计算人工智能模型与认知智能重点实验室) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) Pazhou Lab(琶洲实验室) Engineering Research Center of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human(教育部健康智能感知与平行数字人工程研究中心)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出多级自适应去混淆方法,通过消除多模态数据中的类间和样本特定混淆,提升多模态模型的分类可靠性。

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23126 2025-12-01 econ.GN q-fin.EC 78%

Does cross-modal discounting generalize to non-WEIRD cultures? A comparison of the USA and Japan

跨模态折扣是否能推广到非WEIRD文化?对美国和日本的比较

Shohei Yamamoto, Rebecca McDonald, Daniel Read

专题命中 其他多模态 :cross-modal(title,abstract)

AI总结 该研究比较了美国和日本,发现跨模选择在不同文化中均表现出更大的耐心,且效应在被试间实验中更显著,支持跨模效应的普遍性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22174 2025-12-01 cs.LO math.LO 78%

Nested Sequents for Intuitionistic Multi-Modal Logics: Cut-Elimination and Lyndon Interpolation

嵌套序列为直觉多模逻辑:切消和Lyndon插值

Tim S. Lyon

专题命中 其他多模态 :multi-modal(title,abstract)

AI总结 本文提出了一种单结论嵌套序列表演算,用于直觉多模逻辑,并证明了切消性和Lyndon插值性质。

Comments in review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09464 2025-12-01 cs.GR cs.CV 74%

Hybrid Rendering for Multimodal Autonomous Driving: Merging Neural and Physics-Based Simulation

多模态自动驾驶的混合渲染:融合神经网络与基于物理的模拟

Máté Tóth, Péter Kovács, Réka Bencses, Zoltán Bendefy, Zoltán Hortsin, Balázs Teréki, Tamás Matuszka

机构 * aiMotive

专题命中 其他多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出混合渲染方法,结合神经网络与基于物理的模拟,提升自动驾驶模拟中新颖视角合成质量及实时渲染效率。

详情

展开后加载摘要…

URL PDF HTML 收藏