arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-07 至 2025-10-07 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 8 篇

2507.09747 2025-10-07 cs.NE 82%

BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings

Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei, Quanying Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22962 2025-10-07 cs.LG cond-mat.mtrl-sci physics.chem-ph 78%

Multimodal machine learning with large language embedding model for polymer property prediction

Tianren Zhang, Dai-Bei Yang

机构 * Department of Materials Science and Engineering, University of Delaware, Newark, Delaware 19716, United States(材料科学与工程系,德雷克塞尔大学) Department of Chemistry, University of Pennsylvania, Philadelphia, Pennsylvania 19104, United States(化学系,宾夕法尼亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref Chem. Mater. 2025, 37, 7002-7013

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16357 2025-10-07 cs.CV 77%

Law of Vision Representation in MLLMs

Shijia Yang, Bohan Zhai, Quanzeng You, Jianbo Yuan, Hongxia Yang, Chenfeng Xu

机构 * Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

Comments The code is available at https://github.com/bronyayang/Law_of_Vision_Representation_in_MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10666 2025-10-07 astro-ph.SR astro-ph.GA astro-ph.IM 75%

Machine-learning inference of stellar properties using integrated photometric and spectroscopic data

Ilay Kamai, Alex M. Bronstein, Hagai B. Perets

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract)

Comments Accepted to ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03455 2025-10-07 cs.CV 70%

PEaRL: Pathway-Enhanced Representation Learning for Gene and Pathway Expression Prediction from Histology

Sejuti Majumder, Saarthak Kapse, Moinak Bhattacharya, Xuan Xu, Alisa Yurovsky, Prateek Prasanna

机构 * Department of Biomedical Informatics, Stony Brook University, NY, USA(生物医学信息学系,石溪大学,纽约,美国) Department of Computer Science, Stony Brook University, NY, USA(计算机科学系,石溪大学,纽约,美国)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04417 2025-10-07 cs.LG cs.AI cs.CL cs.CV cs.IT math.IT 67%

Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions

Wenyuan Zhao, Adithya Balachandran, Chao Tian, Paul Pu Liang

机构 * Texas A&M University(德克萨斯大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20752 2025-10-07 cs.CV cs.AI 62%

Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models

Huajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院自动化研究所人工智能学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 51 pages, 23 figures, NeurIPS'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04196 2025-10-07 cs.AI cs.LG 57%

COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability

Yizhuo Ding, Mingkang Chen, Qiuhua Liu, Fenghua Weng, Wanying Qu, Yue Yang, Yugang Jiang, Zuxuan Wu, Yanwei Fu, Wenqi Shao

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shenzhen University(深圳大学) ShanghaiTech University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏