arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-25 至 2025-08-25 共收录 12 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 12 篇

2508.16300 2025-08-25 cs.CV cs.AI 88%

A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension

Mohammad Zia Ur Rehman, Devraj Raghuvanshi, Umang Jain, Shubhi Bansal, Nagendra Kumar

机构 * Indian Institute of Technology Indore(印度印度理工学院Indore) Brown University(布朗大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Published in Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15653 2025-08-25 cs.CV 85%

MapKD: Unlocking Prior Knowledge with Cross-Modal Distillation for Efficient Online HD Map Construction

Ziyang Yan, Ruikai Li, Zhiyong Cui, Bohan Li, Han Jiang, Yilong Ren, Aoyong Li, Zhenning Li, Sijia Wen, Haiyang Yu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16147 2025-08-25 cs.IR 82%

Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction

Ao Zhou, Mingsheng Tu, Luping Wang, Tenghao Sun, Zifeng Cheng, Yafeng Yin, Zhiwei Jiang, Qing Gu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16230 2025-08-25 cs.CV cs.AI 81%

FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing

Jiahao Chen, Zhiyong Ma, Wenbiao Du, Qingyuan Chuai

机构 * Cao Tu Li(Guangzhou) Technology Co., Ltd, China(广东曹图利技术有限公司) South China University of Technology, China(华南理工大学)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01068 2025-08-25 cs.CL cs.AI 81%

Multimodal Transformers are Hierarchical Modal-wise Heterogeneous Graphs

Yijie Jin, Junjie Peng, Xuanchao Lin, Haochen Yuan, Lan Wang, Cangzhi Zheng

机构 * School of Computer Engineering and Science, Shanghai University(上海大学计算机工程与科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Journal ref https://aclanthology.org/2025.acl-long.109/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16408 2025-08-25 cs.CV 79%

SAMFusion: Sensor-Adaptive Multimodal Fusion for 3D Object Detection in Adverse Weather

Edoardo Palladin, Roland Dietze, Praveen Narayanan, Mario Bijelic, Felix Heide

机构 * Torc Robotics(Torc机器人公司) University of Stuttgart(斯图加特大学) Princeton University(普林斯顿大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16000 2025-08-25 eess.IV cs.CV cs.LG 79%

Cross-Attention Multimodal Fusion for Breast Cancer Diagnosis: Integrating Mammography and Clinical Data with Explainability

Muhaisin Tiyumba Nantogmah, Abdul-Barik Alhassan, Salamudeen Alhassan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15852 2025-08-25 cs.LG cs.CL 79%

PGF-Net: A Progressive Gated-Fusion Framework for Efficient Multimodal Sentiment Analysis

Bin Wen, Tien-Ping Tan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16054 2025-08-25 cs.AI cs.CL 79%

Generative Foundation Model for Structured and Unstructured Electronic Health Records

Sonish Sivarajkumar, Hang Zhang, Yuelyu Ji, Maneesh Bilalpur, Xizhi Wu, Chenyu Li, Min Gu Kwak, Shyam Visweswaran, Yanshan Wang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16008 2025-08-25 cs.RO 71%

Self-Aligning EPM Connector: A Versatile Solution for Adaptive and Multi-Modal Interfaces

Bingchao Wang, Adam A. Stokes

机构 * Soft Systems Group, The School of Engineering, The University of Edinburgh(软系统组、工程学院、爱丁堡大学)

专题命中 多模态训练与对齐 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15977 2025-08-25 cs.CL 57%

Dancing with Deer: A Constructional Perspective on MWEs in the Era of LLMs

Claire Bonial, Julia Bonn, Harish Tayyar Madabushi

机构 * U.S. Army Research Lab(美国陆军研究实验室) University of Colorado Boulder(科罗拉多大学博尔德分校) University of Bath(巴斯大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL

Comments Chapter in Phraseology and Multiword Expressions, Language Science Press (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20182 2025-08-25 cs.LG cs.AI q-bio.QM 57%

Chemical Language Model Linker: blending text and molecules with modular adapters

Yifan Deng, Spencer S. Ericksen, Anthony Gitter

机构 * Department of Computer Sciences, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系) Drug Development Core, Small Molecule Screening Facility, University of Wisconsin Carbone Cancer Center, University of Wisconsin-Madison(威斯康星大学卡本纳癌症中心药物开发核心、小分子筛选设施) Department of Biostatistics and Medical Informatics, University of Wisconsin-Madison(威斯康星大学麦迪逊分校生物统计学与医学信息学系)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments 63 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏