arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-22 至 2025-09-22 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 11 篇

2509.15578 2025-09-22 cs.CV cs.AI 84%

Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality Fusion

Shanghong Li, Chiam Wen Qi Ruth, Hong Xu, Fang Liu

机构 * Singapore University of Social Sciences(新加坡社会科学研究大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16149 2025-09-22 cs.CV 83%

Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models

Renjie Pi, Kehao Miao, Li Peihang, Runtao Liu, Jiahui Gao, Jipeng Zhang, Xiaofang Zhou

机构 * HKUST(香港科技大学) HKU(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16017 2025-09-22 cs.CV 83%

DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching

Meng Yang, Fan Fan, Zizhuo Li, Songchu Deng, Yong Ma, Jiayi Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15852 2025-09-22 cs.MM 79%

Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction

Yueheng Jiang, Peng Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15277 2025-09-22 cs.MM cs.LG 79%

Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction

Qin Chao, Eunsoo Kim, Boyang Li

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡) Business School, University of Seoul, Republic of Korea(商学学院,首尔大学,韩国) Alibaba Group and the Alibaba-NTU Joint Research Institute, Singapore(阿里巴巴集团及阿里巴巴-南洋理工大学联合研究机构,新加坡)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19668 2025-09-22 eess.SP cs.AI cs.CL cs.LG 76%

SuPreME: A Supervised Pre-training Framework for Multimodal ECG Representation Learning

Mingsheng Cai, Jiuming Jiang, Wenhao Huang, Che Liu, Rossella Arcucci

机构 * The University of Edinburgh(爱丁堡大学) Imperial College London(帝国理工学院) Shenzhen Yinwang Intelligent Technology Co., Ltd(深圳英伟达智能技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL、cs.AI

Comments Findings of The 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15667 2025-09-22 cs.CL cs.SD eess.AS 73%

VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos, Vassilis Katsouros

机构 * Institute for Speech and Language Processing, Athena Research Center, Greece(语音与语言处理研究所,亚特兰蒂斯研究中心,希腊)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14067 2025-09-22 cs.CV cs.AI 73%

VLA-Mark: A cross modal watermark for large vision-language alignment model

Shuliang Liu, Qi Zheng, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) University of Toronto(多伦多大学) Ant Group, Alibaba(蚂蚁集团,阿里巴巴) New York University Shanghai(纽约大学上海分校)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by the main conference, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16170 2025-09-22 cs.CV 70%

UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation

Xiaoqi Zhao, Youwei Pang, Chenyang Yu, Lihe Zhang, Huchuan Lu, Shijian Lu, Georges El Fakhri, Xiaofeng Liu

机构 * Yale University, USA(耶鲁大学) Nanyang Technological University, Singapore(南洋理工大学) Dalian University of Technology, China(大连理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15990 2025-09-22 cs.CV 57%

DAFTED: Decoupled Asymmetric Fusion of Tabular and Echocardiographic Data for Cardiac Hypertension Diagnosis

Jérémie Stym-Popper, Nathan Painchaud, Clément Rambour, Pierre-Yves Courand, Nicolas Thome, Olivier Bernard

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 9 pages, Accepted at MIDL 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18353 2025-09-22 cs.LG 50%

A Data-Driven Review of Remote Sensing-Based Data Fusion in Precision Agriculture from Foundational to Transformer-Based Techniques

Mahdi Saki, Rasool Keshavarz, Daniel Franklin, Mehran Abolhasan, Justin Lipman, Negin Shariati

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 22 pages, 13 figures, 3 tables, Journal

详情

展开后加载摘要…

URL PDF HTML 收藏