arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2601.21006 2026-02-05 physics.plasm-ph 78%

A joint diffusion approach to multi-modal inference in inertial confinement fusion

一种联合扩散方法用于惯性约束聚变中的多模态推断

Michael S. Jones, Justin Kunimune, Daniel Casey, Bogdan Kustowski, Eugene Kur, Kelli Humbird

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出JointDiff方法,通过联合扩散统一正向建模、逆向推断和输出填补,提升惯性约束聚变实验的多模态推断精度与可转移性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00508 2026-02-05 eess.SP 78%

Quadrature Over-the-Air-Computing for Multimodal Dual-Stream Signal Processing

正交空天地计算用于多模双流信号处理

Hyeon Seok Rou, Kengo Ando, Giuseppe Thadeu Freitas de Abreu, David González G

专题命中 多模态生成 :multimodal(title);multi-modal(abstract)

AI总结 Q-OTAC通过利用复信号的同相和正交分量,实现双流同时计算,提升计算效率,适用于多模B5G应用。

Comments Accepted at the IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02927 2026-02-04 stat.ML cs.LG 78%

Training-Free Self-Correction for Multimodal Masked Diffusion Models

无需训练的多模态掩码扩散模型自校正

Yidong Ouyang, Panwen Hu, Zhengyan Wan, Zhe Wang, Liyan Xie, Dmitriy Bespalov, Ying Nian Wu, Guang Cheng, Hongyuan Zha, Qiang Sun

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Mohamed bin Zayed University of Artificial Intelligence(莫莫德·本·扎耶德人工智能大学) East China Normal University(华东师范大学) University of Virginia(弗吉尼亚大学) University of Minnesota(明尼苏达大学) Drexel university(德雷塞尔大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Toronto(多伦多大学)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出无需训练的多模态掩码扩散模型自校正方法,通过减少采样步骤提升生成质量,适用于文本到图像和多模态理解任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00641 2026-02-03 stat.ML cs.LG stat.CO 78%

Sampling from multi-modal distributions on Riemannian manifolds with training-free stochastic interpolants

在黎曼流形上采样多模分布的无训练随机插值法

Alain Durmus, Maxence Noble, Thibaut Pellerin

机构 * CMAP, CNRS, Ecole polytechnique(CMAP、法国国家科学研究中心、巴黎高等师范学院)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出了一种无训练的随机插值方法,用于在黎曼流形上高效采样多模分布,通过非平衡动力学和随机插值实现无需训练的高效采样。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22407 2026-02-02 cond-mat.mtrl-sci 78%

Nanoscale mapping of phase-transformation pathways in medium-Mn TRIP steel by multimodal STEM

中等锰TRIP钢中相变路径的纳米级映射:多模态STEM

Marc Raventós-Tato, S. Leila Panahi, Núria Bagués, David Frómeta, Oleg Usoltsev, Núria Cuadrado, Joaquín Otón

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本研究通过多模态STEM技术,实现了中等锰TRIP钢中相变路径的纳米级映射,结合电子衍射与能谱分析,实现了铁、奥氏体和马氏体的相分离与晶格参数细化。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21950 2026-01-30 cs.LG 78%

Embracing Aleatoric Uncertainty in Medical Multimodal Learning with Missing Modalities

在医疗多模态学习中拥抱概率不确定性与缺失模态

Linxiao Gong, Yang Liu, Lianlong Sun, Yulai Bi, Jing Liu, Xiaoguang Zhu

机构 * HKUST (GZ)(香港科技大学) Tongji University(同济大学) University of Rochester(罗切斯特大学) Meta Fudan University(复旦大学) The University of British Columbia(不列颠哥伦比亚大学) University of California, Davis(加州大学戴维斯分校)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出AUM框架,通过建模单模态概率不确定性来应对医疗多模态学习中的缺失模态问题,在死亡预测任务中取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21129 2026-01-30 cs.RO 78%

WheelArm-Sim: A Manipulation and Navigation Combined Multimodal Synthetic Data Generation Simulator for Unified Control in Assistive Robotics

WheelArm-Sim: 一种结合 manipulation 和 navigation 的多模态合成数据生成模拟器,用于辅助机器人中的统一控制

Guangping Liu, Tipu Sultan, Vittorio Di Giorgio, Nick Hawkins, Flavio Esposito, Madi Babaiasl

机构 * Aerospace and Mechanical Engineering Department, Saint Louis University(圣路易斯大学航空航天与机械工程系) Computer Science Department, Saint Louis University(圣路易斯大学计算机科学系)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 WheelArm-Sim 是一种用于辅助机器人统一控制的多模态合成数据生成模拟器,通过集成轮椅和机械臂控制,为数据驱动的机器学习模型提供支持。

Comments Accepted to IEEE International Symposium on Medical Robotics (ISMR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20406 2026-01-27 cs.RO cs.LG 78%

PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning

PointMapPolicy: 结构化点云处理用于多模态模仿学习

Xiaogang Jia, Qian Wang, Anrui Wang, Han A. Wang, Balázs Gyenes, Emiliyan Gospodinov, Xinkai Jiang, Ge Li, Hongyi Zhou, Weiran Liao, Xi Huang, Maximilian Beck, Moritz Reuss, Rudolf Lioutikov, Gerhard Neumann

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Reality Labs, Meta(Meta现实实验室) Johannes Kepler University Linz(林茨约翰尼斯·开普勒大学)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 PointMapPolicy通过结构化点云处理提升多模态模仿学习的精度与泛化能力,利用xLSTM融合点云与RGB数据,在RoboCasa和CALVIN基准中取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02900 2026-01-22 cs.CV cs.AI cs.GR cs.HC cs.MM 78%

Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions

推进说话头生成:多模态方法、数据集、评估指标和损失函数的全面综述

Vineet Kumar Rakesh, Soumya Mazumdar, Research Pratim Maity, Sarbajit Pal, Amitabha Das, Tapas Samanta

机构 * Engineering Sciences, Homi Bhabha National Institute Training School Complex(工程科学,霍米·巴赫瓦全国研究所培训学校) Computer and Informatics Group, VECC 1/AF(计算机与信息组,VECC 1/AF)

专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.AI、cs.MM

AI总结 本文全面综述了说话头生成的多模态方法、数据集、评估指标和损失函数,探讨了技术挑战与未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12381 2026-01-21 q-bio.QM q-bio.BM q-bio.GN 78%

Multimodal Spatial Omics: From Data Acquisition to Computational Integration

多模态空间组学:从数据获取到计算整合

Esra Busra Isik, Yusuf Hakan Usta, Haozhe Liu, Maryam Riazi, William Roach, Hongpeng Zhou, Magnus Rattray, Sokratia Georgaka

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文综述了多模态空间组学数据整合的计算方法,涵盖从概率模型到深度学习的多种算法原理,以整合不同分子层和图像数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12277 2026-01-21 cs.RO 78%

An Efficient and Multi-Modal Navigation System with One-Step World Model

一种高效且多模态的导航系统与一步世界模型

Wangtian Shen, Ziyang Meng, Jinming Ma, Mingliang Zhou, Diyun Xiang

机构 * Tsinghua University(清华大学) Xiaomi (China)(小米(中国))

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出了一种高效多模态导航系统,通过一步世界模型和3D U-Net骨干网络,提升导航效率和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11540 2026-01-21 cs.HC 78%

Exploring General-Purpose Autonomous Multimodal Agents for Pathology Report Generation

探索通用型自主多模态代理用于病理报告生成

Marc Aubreville, Taryn A. Donovan, Christof A. Bertram

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 研究探索通用型自主多模态代理在病理报告生成中的应用,发现其在无额外信息时诊断准确率较低,但提供形态学描述时表现更佳,但仍不及人类专家。

Comments 6 pages, 1 figure, accepted paper for BVM 2026

Journal ref BVM 2026, https://bvm-conf.org

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02766 2026-01-07 cs.RO cs.AR 78%

Advancing Assistive Robotics: Multi-Modal Navigation and Biophysical Monitoring for Next-Generation Wheelchairs

推动辅助机器人:多模态导航与生物物理监测用于下一代轮椅

Md. Anowar Hossain, Mohd. Ehsanul Hoque

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出了一种多模态轮椅控制系统,结合多种输入接口与生物物理监测,提升患者独立性并实现护理人员实时监督。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01475 2026-01-06 cs.LG 78%

Multi-Subspace Multi-Modal Modeling for Diffusion Models: Estimation, Convergence and Mixture of Experts

多子空间多模态建模用于扩散模型:估计、收敛与专家混合

Ruofeng Yang, Yongcan Li, Bo Jiang, Cheng Chen, Shuai Li

机构 * Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出基于多子空间多模态建模的扩散模型,通过引入低秩混合高斯混合结构,有效捕捉多模态信息并提升生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04000 2026-01-01 cs.IT cs.LG math.IT 78%

Distributed Information Bottleneck Theory for Multi-Modal Task-Aware Semantic Communication

多模态任务感知语义通信的分布式信息瓶颈理论

Yujie Zhou, Cheng Peng, Rulong Wang, Yong Xiao, Yingyu Li, Guangming Shi, Ping Zhang

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息与通信学院,华中科技大学) Peng Cheng Laboratory(鹏城实验室) Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔)) School of Mechanical Engineering and Electronic Information, China University of Geosciences(机械工程与电子信息学院,中国地质大学) State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出了一种多模态任务感知语义通信的分布式信息瓶颈框架,通过量化模态贡献来优化资源利用,提升通信效率与任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19983 2025-12-24 cs.IR 78%

IGDMRec: Behavior Conditioned Item Graph Diffusion for Multimodal Recommendation

IGDMRec: 基于行为的物品图扩散用于多模态推荐

Ziyuan Guo, Jie Guo, Zhenghao Chen, Bin Song, Fei Richard Yu

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 IGDMRec通过行为条件图扩散和条件去噪网络,利用用户行为信息去噪多模态推荐的语义物品图,提升推荐性能。

Comments 12 pages, 6 figures. This paper has been accepted for publication in IEEE Transactions on Multimedia. The final published version will be available via IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04359 2025-12-15 eess.SP 78%

Efficient Domain Generalization in Wireless Networks with Scarce Multi-Modal Data

在稀缺多模态数据下实现无线网络高效领域泛化

Minsu Kim, Walid Saad, Doru Calin

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出了一种两阶段学习框架,通过基于物理的损失函数和协作域适应方法,在稀缺多模态数据下提升无线网络的领域泛化性能。

Comments Submitted to IEEE TWC

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15365 2025-12-12 eess.SY cs.LG cs.MA cs.SY 78%

TranSimHub:A Unified Air-Ground Simulation Platform for Multi-Modal Perception and Decision-Making

TranSimHub:一个多模态感知与决策的空地协同仿真平台

Maonan Wang, Yirong Chen, Yuxin Cai, Aoyu Pang, Yuejiao Xie, Zian Ma, Chengcheng Xu, Kemou Jiang, Ding Wang, Laurent Roullet, Chung Shue Chen, Zhiyong Cui, Yuheng Kan, Michael Lepech, Man-On Pun

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shanghai AI Laboratory(上海人工智能实验室) Stanford University(斯坦福大学) Nanyang Technological University(南洋理工大学) SenseTime Group Ltd(商汤科技有限公司) Beihang University(北京航空航天大学) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 TranSimHub是一个用于空地协同智能的统一仿真平台,支持多模态感知与决策研究,提供同步渲染和可控场景编辑功能。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03375 2025-12-04 cs.LG 78%

MAGE-ID: A Multimodal Generative Framework for Intrusion Detection Systems

MAGE-ID:一种多模态生成框架用于入侵检测系统

Mahdi Arab Loodaricheh, Mohammad Hossein Manshaei, Anita Raja

机构 * Department of Computer Science, Hunter College and The Graduate Center, City University of New York(计算机科学系,亨特学院和研究生中心,纽约市立大学)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 MAGE-ID通过多模态生成框架提升入侵检测系统的数据增强效果,实现更平衡和一致的多模态合成,显著提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12770 2025-12-01 cs.LG cs.CE 78%

MolEdit: Knowledge Editing for Multimodal Molecule Language Models

MolEdit: 多模态分子语言模型的知识编辑

Zhenyu Lei, Patrick Soga, Yaochen Zhu, Yinhan He, Yushun Dong, Jundong Li

机构 * University of Virginia(弗吉尼亚大学) Florida State University(佛罗里达州立大学)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 MolEdit通过多专家知识适配器和专家意识编辑切换器,提升多模态分子语言模型的编辑可靠性与局部性,实现分子与描述词的高效互转。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14631 2025-11-12 cs.DB 78%

Towards a Multimodal Stream Processing System

Uélison Jean Lopes dos Santos, Alessandro Ferri, Szilard Nistor, Riccardo Tommasini, Carsten Binnig, Manisha Luthra

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01374 2025-11-04 cs.LG 78%

Learning Intractable Multimodal Policies with Reparameterization and Diversity Regularization

Ziqi Wang, Jiashun Liu, Ling Pan

机构 * Hong Kong University of Science and Technology(香港理工大学)

专题命中 多模态生成 :multimodal(title,abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01074 2025-10-31 cs.LG 78%

Omni-Mol: Multitask Molecular Model for Any-to-any Modalities

Chengxin Hu, Hao Li, Yihe Yuan, Zezheng Song, Chenyang Zhao, Haixin Wang

机构 * National University of Singapore(新加坡国立大学) University of Maryland, College Park(马里兰大学 College Park 分校) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 多模态生成 :any-to-any(title);multimodal(abstract)

Comments 44 pages, 9 figures, 13 tables, paper accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20637 2025-10-24 cs.LG 78%

Large Multimodal Models-Empowered Task-Oriented Autonomous Communications: Design Methodology and Implementation Challenges

Hyun Jong Yang, Hyunsoo Kim, Hyeonho Noh, Seungnyun Kim, Byonghyo Shim

机构 * Seoul National University(首尔国立大学) Hanbat National University(翰baum国立大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20595 2025-10-24 stat.ML cs.LG 78%

Diffusion Autoencoders with Perceivers for Long, Irregular and Multimodal Astronomical Sequences

Yunyi Shen, Alexander Gagliano

机构 * EECS, MIT(麻省理工学院电子工程与计算机科学系) IAIFI, MIT(麻省理工学院天文研究所) CfA, Harvard(哈佛大学天文台)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00565 2025-10-24 stat.CO cs.LG math.ST stat.TH 78%

Sampling from multi-modal distributions with polynomial query complexity in fixed dimension via reverse diffusion

Adrien Vacher, Omar Chehab, Anna Korba

机构 * CREST, ENSAE Institut Polytechnique de Paris(CREST,ENSAE 巴黎高等理工学院)

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18879 2025-10-23 cs.HC 78%

FIRETWIN: Digital Twin Advancing Multi-Modal Sensing, Interactive Analytics for Wildfire Response

Mayamin Hamid Raha, Ali Reza Tavakkoli, Chris Webb, Mobin Habibpour, Janice Coen, Eric Rowell, Fatemeh Afghah

专题命中 多模态生成 :multi-modal(title);multimodal(abstract)

Comments 8 pages, 6 figures, accepted in IEEE International Workshop on Computer-Aided Modeling and Design of Communication Links and Networks (CAMAD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10914 2025-10-14 eess.SY cs.SY 78%

Optimal Multi-Modal Transportation and Electric Power Flow: The Value of Coordinated Dynamic Operation

Jiajie Qiu, Dakota Thompson, Kamal Youcef-Toumi, Amro M. Farid

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 31 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24840 2025-10-13 cs.LG cs.CE 78%

Cell2Text: Multimodal LLM for Generating Single-Cell Descriptions from RNA-Seq Data

Oussama Kharouiche, Aris Markogiannakis, Xiao Fei, Michail Chatzianastasis, Michalis Vazirgiannis

机构 * École Polytechnique, IP Paris(巴黎理工学院,IP巴黎)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21874 2025-09-29 cs.LG 78%

Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models

Yifei Peng, Yaoli Liu, Enbo Xia, Yu Jin, Wang-Zhou Dai, Zhong Ren, Yao-Xiang Ding, Kun Zhou

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) National Key Laboratory for Novel Software Technology(新型软件技术国家实验室)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏