arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-17 至 2025-12-17 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 9 篇

2508.00969 2025-12-17 cs.LG cs.AI 87%

Masked Omics Modeling for Multimodal Representation Learning across Histopathology and Molecular Profiles

掩码组学建模用于病理学与分子特征的多模态表示学习

Lucas Robinet, Ahmad Berjaoui, Elizabeth Cohen-Jonathan Moyal

机构 * Oncopole(奥恩波尔) IRT Saint Exupéry(国际研究与技术圣埃克苏佩里) INSERM Cancer Research Center of Toulouse(里沃利癌症研究中心) Toulouse(图卢兹)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);any-to-any(abstract);multimodal foundation model(abstract)

AI总结 MORPHEUS通过整合病理学图像和多组学数据,提出了一种多模态预训练策略,以提升癌症研究中的跨模态表示学习能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11160 2025-12-17 cs.CV 79%

A Unified Framework with Multimodal Fine-tuning for Remote Sensing Semantic Segmentation

多模态微调的统一框架用于遥感语义分割

Xianping Ma, Xiaokang Zhang, Man-On Pun, Bo Huang

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)科学与工程学院) School of Information Science and Engineering, Wuhan University of Science and Technology(武汉科技大学信息科学与工程学院) Department of Geography, The University of Hong Kong(香港大学地理系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种多模态微调的统一框架,用于改进遥感语义分割的性能,通过引入新的MFNet和DFM模块,显著提升了多模态数据的分割效果。

Comments 15 pages, 11 figures

Journal ref IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1-15, 2025, Art no. 5405015

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22780 2025-12-17 cs.LG physics.geo-ph 78%

Multimodal Atmospheric Super-Resolution With Deep Generative Models

多模态大气超分辨率与深度生成模型

Dibyajyoti Chakraborty, Haiwen Guan, Jason Stock, Troy Arcomano, Guido Cervone, Romit Maulik

机构 * Information Sciences and Technology(信息科学与技术系) The Pennsylvania State University(宾夕法尼亚州立大学) Environmental Science Division(环境科学部) Argonne National Laboratory(阿贡国家实验室) Department of Geography(地理系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出利用基于分数的扩散模型进行多模态大气超分辨率,通过结合稀疏观测数据提升高维状态恢复精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13177 2025-12-17 cs.CV cs.RO 70%

MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion

MMDrive: 通过多表示融合超越视觉的交互场景理解

Minghui Hou, Wei-Hsing Huang, Shaofeng Liang, Daizong Liu, Tai-Hao Wen, Gang Wang, Runwei Guan, Weiping Ding

机构 * organization= College of Computer Science Technology, Jilin University , city= Changchun , country= China organization= Georgia Institute of Technology , city= Atlanta , country= USA organization= Qingdao Institute of Software, College of Computer Science Technology, China University of Petroleum (East China) , city= Qingdao , country= China organization= Institute for Math \& AI, Wuhan University , city= Wuhan , country= China organization= University of Michigan, Ann Arbor , country= USA organization= Thrust of Artificial Intelligence, Hong Kong University of Science organization= School of Artificial Intelligence Computer Science, Nantong University , city= Nantong , country= China

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 MMDrive通过融合占用图、LiDAR点云和文本描述,实现超越视觉的三维场景理解,提升自动驾驶的多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14026 2025-12-17 cs.CV 57%

Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers

通过打破跨表格障碍释放图像-表格自监督学习的潜力

Yibing Fu, Yunpeng Zhao, Zhitao Zeng, Cheng Chen, Yueming Jin

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出CITab框架,通过跨表格学习提升医学图像与表格数据的多模态自监督学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19981 2025-12-17 cs.CV 57%

FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking

FutrTrack: 一种基于相机与激光雷达融合的变压器用于三维多目标跟踪

Martha Teiko Teye, Ori Maoz, Matthias Rottmann

机构 * University of Wuppertal(乌珀塔尔大学) Institute of Computer Science, Osnabrück University(计算机科学研究所,奥斯纳布鲁克大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 FutrTrack通过融合相机与激光雷达数据,利用基于变压器的跟踪框架,在多目标跟踪中实现了高效且鲁棒的跟踪性能。

Comments Accepted to VISAPP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14608 2025-12-17 eess.SP 50%

Fusion of Cellular ISAC and Passive RF Sensing for UAV Detection and Tracking

细胞ISAC与被动RF传感的融合用于UAV检测与跟踪

Cole Dickerson, Sean Kearney, Sultan Manjur, Ismail Guvenc, Sevgi Gurbuz, Ali Gurbuz, Ozgur Ozdemir, Mihail Sichitiu

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 本文提出了一种融合被动RF传感和雷达的无人机检测与跟踪系统,通过卡尔曼滤波器整合异步观测结果,提升跟踪精度和覆盖范围。

Comments Accepted for publication at the 2025 IEEE Asilomar Conference on Signals, Systems, and Computers, Session: UAV Intrusion Detection Using Mobile Communications Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11065 2025-12-17 cs.HC 50%

Immutable Explainability: Fuzzy Logic and Blockchain for Verifiable Affective AI

不可变的可解释性:模糊逻辑与区块链用于可验证的情感人工智能

Marcelo Fransoy, Alejandro Hossian, Hernán Merlino

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出不可变可解释性架构,结合模糊逻辑与区块链,实现情感AI的透明决策和可信审计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14019 2025-12-17 cs.LG q-bio.QM 50%

EXAONE Path 2.5: Pathology Foundation Model with Multi-Omics Alignment

EXAONE Path 2.5:多组学对齐的病理基础模型

Juseung Yun, Sunwoo Yu, Sumin Ha, Jonghyun Kim, Janghyeon Lee, Jongseong Jang, Soonyoung Lee

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 EXAONE Path 2.5通过多组学对齐构建病理基础模型,实现更全面的肿瘤生物学建模,展现高效率和适应性,推动精准肿瘤学发展。

详情

展开后加载摘要…

URL PDF HTML 收藏