arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-27 至 2026-01-27 共收录 24 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 24 篇

2601.17986 2026-01-27 cs.LG 85%

Federated learning for unpaired multimodal data through a homogeneous transformer model

通过同质Transformer模型实现无配对多模态数据的联邦学习

Anders Eklund

机构 * Department of Biomedical Engineering, Linköping University, Sweden(_linköping大学生物医学工程系) Department of Computer and Information Science, Linköping University, Sweden(_linköping大学计算机与信息科学系) Center for Medical Image Science and Visualization (CMIV), Linköping University, Sweden(_linköping大学医学影像科学与可视化中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);multimodal foundation model(abstract)

AI总结 本文提出通过同质Transformer模型实现联邦学习,解决无配对多模态数据的训练问题,通过公共锚点对齐和子空间稳定化微调方法,在不传输私有数据的情况下实现全局模型统一表示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19327 2026-01-27 cs.CV cs.AI 84%

DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding

DeepInsert: 早期层绕过以实现高效且高性能的多模态理解

Moulik Choraria, Xinbo Wu, Akhil Bhimaraju, Nitesh Sekhar, Yue Wu, Xu Zhang, Prateek Singhal, Lav R. Varshney

机构 * UIUC(伊利诺伊大学) Amazon(亚马逊) Capital One Apple(苹果) Stony Brook University(石溪大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 DeepInsert通过将多模态令牌插入模型中间层以绕过早期层,从而在降低训练和推理成本的同时保持或提升多模态语言模型的性能。

Comments To be presented at EACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15139 2026-01-27 cs.CV 83%

Unified Cross-Modal Attention-Mixer Based Structural-Functional Connectomics Fusion for Neuropsychiatric Disorder Diagnosis

统一的跨模态注意力-混合器基于结构-功能连接组融合的神经精神疾病诊断

Badhan Mazumder, Lei Wu, Vince D. Calhoun, Dong Hye Ye

机构 * Department of Computer Science, Georgia State University(计算机科学系,佐治亚州立大学) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, and Emory University(跨机构神经影像与数据科学转化研究中心(TReNDS),佐治亚州立大学、佐治亚理工学院和埃默里大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出ConneX方法,通过统一的跨模态注意力和MLP-Mixer实现结构-功能连接组的多模态融合,提升神经精神疾病诊断性能。

Comments Published in the Proceedings of the 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2025). IEEE Xplore. DOI: 10.1109/EMBC58623.2025.11254194

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17809 2026-01-27 cs.IT math.IT 82%

A Multi-Modal Fusion Platform for Joint Environment Sensing and Channel Sounding in Highly Dynamic Scenarios

一种多模态融合平台,用于高动态场景中的联合环境感知与信道探测

Xuejian Zhang, Ruisi He, Mi Yang, Zhengyu Zhang, Ziyi Qi

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)

AI总结 本文提出一种多模态融合平台,用于高动态场景中联合环境感知与信道探测,支持多频段多天线测量和高精度环境感知。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15304 2026-01-27 cs.IR 82%

MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation

MLLMRec: 基于图细化的多模态推荐偏好推理范式

Yuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan, Wanyu Chen, Fei Cai, Honghui Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

AI总结 MLLMRec通过图细化和多模态大语言模型提升多模态推荐的用户偏好推理与物品表示学习准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18589 2026-01-27 cs.CV cs.MM 81%

AGSP-DSA: An Adaptive Graph Signal Processing Framework for Robust Multimodal Fusion with Dynamic Semantic Alignment

AGSP-DSA: 一种用于鲁棒多模态融合的自适应图信号处理框架,具有动态语义对齐

KV Karthikeya, Ashok Kumar Das, Shantanu Pal, Vivekananda Bhat K, Arun Sekar Rajasekaran

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 AGSP-DSA通过自适应图信号处理和动态语义对齐,实现鲁棒的多模态融合,提升情感分析和多媒体分类的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21885 2026-01-27 cs.CV cs.MM cs.RO 81%

Integrating Multi-Modal Sensors: A Review of Fusion Techniques for Intelligent Vehicles

多模态传感器整合:智能车辆融合技术综述

Chuheng Wei, Ziye Qin, Ziyan Zhang, Guoyuan Wu, Matthew J. Barth

机构 * College of Engineering, Center for Environmental Research and Technology, University of California at Riverside(工程学院、环境研究与技术中心、加州大学河滨分校) School of Transportation and Logistics, Southwest Jiaotong University(交通运输与物流学院、西南交通大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.MM

AI总结 本文综述了多传感器融合技术在自动驾驶中的应用,分析了深度学习方法、多模态数据集及新兴趋势,强调了其提升系统适应性和鲁棒性的潜力。

Comments Accepted by IEEE IV 2025

Journal ref Proceedings of the 2025 IEEE Intelligent Vehicles Symposium (IV), pp. 1817-1824, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11740 2026-01-27 cs.LG cs.CV 79%

Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent

通过模态解耦梯度下降缓解MLLM指令微调中的视觉知识遗忘

Junda Wu, Yuxin Xiong, Xintong Li, Yu Xia, Ruoyu Wang, Yu Wang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Lina Yao, Jingbo Shang, Julian McAuley

机构 * UC San Diego(加州大学圣迭戈分校) University of New South Wales(新南威尔士大学) Adobe Research(Adobe研究)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出模态解耦梯度下降方法,通过有效秩量化视觉退化并缓解过度压缩,保留预训练视觉知识并提升任务适应能力。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18683 2026-01-27 stat.ME astro-ph.IM cs.LG 78%

Learned harmonic mean estimation of the marginal likelihood for multimodal posteriors with flow matching

基于流匹配的连续归一化流用于多模后验的学得谐均估计

Alicja Polanska, Jason D. McEwen

机构 * Mullard Space Science Laboratory, University College London(穆尔空间科学实验室,伦敦大学学院) Alan Turing Institute(艾伦·图灵研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出基于流匹配的连续归一化流,用于高效估计多模后验的边际似然,解决了传统方法在处理复杂后验时的局限性。

Comments Submitted to 44th International Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05840 2026-01-27 cs.LG 78%

Multimodal Trajectory Representation Learning for Travel Time Estimation

多模态轨迹表示学习用于旅行时间估计

Zhi Liu, Xuyuan Hu, Xiao Han, Zhehao Dai, Zhaolin Deng, Guojiang Shen, Xiangjie Kong

机构 * Zhejiang University of Technology(浙江工业大学) Zhejiang University of Technology College of Computer Science(浙江工业大学计算机科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出多模态动态轨迹整合框架,通过整合GPS、网格轨迹和道路网络约束,提升旅行时间估计的性能和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12869 2026-01-27 cs.LG cs.AI cs.DC cs.IT cs.MA math.IT 70%

On the Fundamental Limits of LLMs at Scale

在大规模下的大语言模型根本限制

Muhammad Ahmed Mohsin, Muhammad Umer, Ahsan Bilal, Zeeshan Memon, Muhammad Ibtsaam Qadir, Sagnik Bhattacharya, Hassan Rizwan, Abhiram R. Gorle, Maahe Zehra Kazmi, Nukhba Amir, Ali Subhan, Muhammad Usman Rafique, Zihao He, Pulkit Mehta, Muhammad Ali Jamshed, John M. Cioffi

机构 * Stanford University(斯坦福大学) The University of Oklahoma(俄克拉荷马大学) Emory University(埃默里大学) Purdue University(普渡大学) UC Riverside(加州大学河滨分校) UC Berkeley(加州大学伯克利分校) Khyber Medical University(克希伯医学大学) Universtat Pompeu Fabra(庞培法华大学) Zoox(Zoox公司) Meta Google DeepMind(谷歌DeepMind) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Glasgow(格拉斯哥大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文探讨了大规模大语言模型的根本限制,提出统一框架分析计算、信息和学习的基础限制,并提供缓解方法。

Comments Submitted to TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09828 2026-01-27 cs.CV cs.LG cs.RO 70%

DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception

DGFusion:基于深度的传感器融合用于鲁棒的语义感知

Tim Broedermannn, Christos Sakaridis, Luigi Piccinelli, Wim Abbeloos, Luc Van Gool

机构 * Computer Vision Laboratory, ETH Zurich(计算机视觉实验室,苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 DGFusion通过整合深度信息提升多模态传感器融合,实现自动驾驶中的鲁棒语义感知。

Comments Code and models are available at https://github.com/timbroed/DGFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23118 2026-01-27 cs.CV 70%

Quantizing Space and Time: Fusing Time Series and Images for Earth Observation

量化空间与时间:融合时间序列和图像用于地球观测

Gianfranco Basile, Johannes Jakubik, Benedikt Blumenstiel, Thomas Brunschwiler, Juan Bernabe Moreno

机构 * IBM Research Europe(IBM欧洲研究院) ETH Zürich(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种任务无关的多模态融合框架,通过时间序列和图像的统一表示空间提升地球观测任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17073 2026-01-27 cs.LG cs.CV stat.ML 70%

Attention-Based Variational Framework for Joint and Individual Components Learning with Applications in Brain Network Analysis

基于注意力的变分框架用于联合和个体成分学习及其在脑网络分析中的应用

Yifei Zhang, Meimei Liu, Zhengwu Zhang

机构 * Department of Biostatistics, Yale School of Public Health(生物统计学系,耶鲁公共卫生学院) Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工大学) Department of Statistics and Operations Research, University of North Carolina at Chapel Hill(统计学与运筹学系,北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出CM-JIVNet,一种基于注意力的变分框架,用于联合和个体成分学习,以提升脑网络分析中的跨模态重建和行为预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18625 2026-01-27 cs.CV 57%

CONQUER: Context-Aware Representation with Query Enhancement for Text-Based Person Search

基于查询增强的上下文感知表示方法用于基于文本的人检索

Zequn Xie

机构 * Zhejiang University, Hangzhou, China(浙江大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 CONQUER通过增强跨模态对齐和自适应查询细化,提升基于文本的人检索的准确性和鲁棒性。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18065 2026-01-27 cs.CL 57%

Grounded Concreteness: Human-Like Concreteness Sensitivity in Vision-Language Models

grounded concreteness: 人类-like 的 concreteness 敏感性在 vision-language 模型中

Aryan Roy, Zekun Wang, Christopher J. MacLellan

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文研究了视觉语言模型在纯文本提示下对concreteness的敏感性,并发现其在更具体的输入上表现更优,具有更清晰的表示和更符合人类规范的判断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11609 2026-01-27 stat.ML cs.AI cs.LG stat.ME 57%

Towards Interpretable Deep Generative Models via Causal Representation Learning

通过因果表示学习实现可解释的深度生成模型

Gemma E. Moran, Bryon Aragam

机构 * Department of Statistics, Rutgers University(统计学系,罗格斯大学) Booth School of Business, University of Chicago(商学院,芝加哥大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出通过因果表示学习实现可解释的深度生成模型,结合潜变量模型、因果图模型和非参数统计,旨在提升生成模型的可解释性和迁移能力。

Comments Accepted in Journal of the American Statistical Association: Special Issue on AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17342 2026-01-27 cs.CV 57%

STARS: Shared-specific Translation and Alignment for missing-modality Remote Sensing Semantic Segmentation

STARS: 共享-特定翻译与对齐用于缺失模态遥感语义分割

Tong Wang, Xiaodong Zhang, Guanzhou Chen, Jiaqi Wang, Chenxi Liu, Xiaoliang Tan, Wenchao Guo, Xuyang Li, Xuanrui Wang, Zifan Wang

机构 * State Key Laboratory of information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(测绘遥感信息工程国家重点实验室,武汉大学) Electronic Information School, Wuhan University(电子信息学院,武汉大学) Hubei FreerTech Co. Ltd(湖北弗瑞特科技有限公司)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 STARS通过共享-特定翻译与对齐机制,解决多模态遥感中缺失模态带来的语义分割问题,提升鲁棒性和类别识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13390 2026-01-27 cs.CV 57%

Generalizing WiFi Gesture Recognition via Large-Model-Aware Semantic Distillation and Alignment

通过大模型感知的语义蒸馏与对齐实现WiFi手势识别的泛化

Feng-Qi Cui, Yu-Tong Guo, Tianyue Zheng, Jinyang Huang

机构 * Institute of Advanced Technology, University of Science and Technology of China(科学技术大学先进技术研究院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) School of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 GLSDA通过大模型感知的语义蒸馏与对齐提升WiFi手势识别的泛化能力,实现领域内和跨领域任务的高性能识别。

Comments Accepted by IEEE ICPADS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16414 2026-01-27 q-bio.NC cs.CV eess.IV 57%

NeuroKoop: Neural Koopman Fusion of Structural-Functional Connectomes for Identifying Prenatal Drug Exposure in Adolescents

NeuroKoop:神经Koopman融合结构-功能连接组用于识别青少年孕期药物暴露

Badhan Mazumder, Aline Kotoski, Vince D. Calhoun, Dong Hye Ye

机构 * Department of Computer Science, Georgia State University(计算机科学系,佐治亚州立大学) Neuroscience Institute, Georgia State University(神经科学研究所,佐治亚州立大学) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS)(跨机构神经影像与数据科学转化研究中心(TReNDS)) Georgia State University, Georgia Institute of Technology, and Emory University(佐治亚州立大学、佐治亚理工学院和埃默里大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 NeuroKoop通过神经Koopman算子融合结构-功能连接组,提升青少年孕期药物暴露识别的准确性和鲁棒性。

Comments Published in the Proceedings of the 2025 IEEE EMBS International Conference on Biomedical and Health Informatics (BHI). IEEE Xplore. DOI: 10.1109/BHI67747.2025.11269557

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17089 2026-01-27 cs.CV 57%

GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing

GRASP: 基于区域感知的稀疏提示引导方法用于适应遥感图像的多模态大语言模型

Qigan Sun, Chaoning Zhang, Jianwei Zhang, Xudong Wang, Jiehui Xie, Pengcheng Zheng, Haoyu Wang, Sungyoung Lee, Chi-lok Andy Tai, Yang Yang, Heng Tao Shen

机构 * School of Computing, Kyung Hee University(京畿大学计算机学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) College of Computer Science and Information Engineering, Harbin Normal University(哈尔滨师范大学计算机科学与信息工程学院) College of Professional and Continuing Education, The Hong Kong Polytechnic University(香港理工大学专业及继续教育学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 GRASP通过引导区域感知的稀疏提示方法,提升多模态大语言模型在遥感图像任务中的适应性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18326 2026-01-27 cs.LG 50%

Cognitive Fusion of ZC Sequences and Time-Frequency Images for Out-of-Distribution Detection of Drone Signals

基于ZC序列和时频图像的认知融合用于无人机信号的分布外检测

Jie Li, Jing Li, Lu Lv, Zhanyu Ju, Fengkui Gong

机构 * State Key Laboratory of Integrated Services Network(集成服务网络国家重点实验室) Xidian University(西安电子科技大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 本文提出基于ZC序列和时频图像的认知融合算法,用于提升无人机信号分布外检测的性能,实验表明其在识别和检测指标上均有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16516 2026-01-27 cs.LG 50%

Rethinking Large Language Models For Irregular Time Series Classification In Critical Care

重新思考用于危重监护中不规则时间序列分类的大型语言模型

Feixiang Zheng, Yu Wu, Cecilia Mascolo, Ting Dang

机构 * The University of Melbourne, Australia(墨尔本大学) University of Cambridge, UK(剑桥大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文研究了LLMs在不规则ICU时间序列分类中的有效性,发现编码器设计比对齐策略更重要,但LLMs在训练时间和少样本学习中表现欠佳。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19034 2026-01-27 eess.SP 50%

Fast Vortex Beam Alignment for OAM Mode Multiplexing in LOS MIMO Networks

快速涡旋束对齐用于 LOS MIMO 网络中的 OAM 模式复用

Poorya Mollahosseini, Yasaman Ghasempour

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 OrthoVortex 提出了一种快速且精确的 OAM 模式对齐方法,通过跨模式相位识别实现快速对齐,提升 LOS MIMO 网络的容量和信干比。

Comments 13 pages, 12 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏