arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-31 至 2025-10-31 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 4 篇

2504.11331 2025-10-31 cs.CL cs.MM 84%

Dependency Structure Augmented Contextual Scoping Framework for Multimodal Aspect-Based Sentiment Analysis

Hao Liu, Lijun He, Jiaxi Liang, Zhihan Ren, Haixia Bi, Fan Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05696 2025-10-31 cs.CV 77%

MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory

Ana Carolina Condez, Diogo Tavares, João Magalhães

机构 * NOVA LINCS, NOVA School of Science and Technology(NOVA LINCS,NOVA科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Updated version: corresponds to the ACM MM '25 published paper and includes full appendix material

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08645 2025-10-31 cs.LG 75%

When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker Product

Youqi Wu, Jingwei Zhang, Farzan Farnia

机构 * Department of Computer Science & Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26146 2025-10-31 cs.LG 50%

maxVSTAR: Maximally Adaptive Vision-Guided CSI Sensing with Closed-Loop Edge Model Adaptation for Robust Human Activity Recognition

Kexing Liu

专题命中 多模态训练与对齐 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏