arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-20 至 2025-10-20 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 4 篇

2504.19136 2025-10-20 cs.CV cs.AI eess.IV 84%

PAD: Phase-Amplitude Decoupling Fusion for Multi-Modal Land Cover Classification

Huiling Zheng, Xian Zhong, Bin Liu, Yi Xiao, Bihan Wen, Xiaofeng Li

机构 * Sanya Science and Education Innovation Park, Wuhan University of Technology, Sanya 572025, China(武汉理工大学三亚科学教育创新园) School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China(武汉理工大学计算机科学与人工智能学院) Key Laboratory of Ocean Circulation and Waves, Institute of Oceanology, Chinese Academy of Sciences, Qingdao 266071, China(中国科学院海洋循环与波浪重点实验室) Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China(湖北省交通物联网重点实验室) State Key Laboratory of Maritime Technology and Safety, Wuhan University of Technology, Wuhan 430063, China(武汉理工大学航海技术与安全国家重点实验室) College of Oceanography and Ecological Science, Shanghai Ocean University, Shanghai 201306, China(上海海洋大学海洋科学与生态学院) School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China(郑州大学计算机与人工智能学院) Rapid-Rich Object Search Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798(南洋理工大学电子与电气工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15317 2025-10-20 cs.AI 79%

VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data

Tingqiao Xu, Ziru Zeng, Jiayu Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15026 2025-10-20 cs.CV 79%

MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning

Mattia Segu, Marta Tintore Gazulla, Yongqin Xian, Luc Van Gool, Federico Tombari

机构 * Google(谷歌) ETH Zurich(苏黎世联邦理工学院) INSAIT, Sofia University, St. Kliment Ohridski(INSAIT,索菲亚大学,圣克莱孟·奥赫里茨基)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15422 2025-10-20 stat.ML cs.LG 50%

Information Theory in Open-world Machine Learning Foundations, Frameworks, and Future Direction

Lin Wang

机构 * Shenzhen Key Laboratory of Neuropsychiatric Modulation(深圳心理行为调控重点实验室) Shenzhen-Hong Kong Institute of Brain Science(深圳-香港脑科学研究院) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏