arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-31 至 2025-10-31 共收录 39 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 4 篇

2506.05696 2025-10-31 cs.CV 77%

MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory

Ana Carolina Condez, Diogo Tavares, João Magalhães

机构 * NOVA LINCS, NOVA School of Science and Technology(NOVA LINCS,NOVA科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Updated version: corresponds to the ACM MM '25 published paper and includes full appendix material

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08645 2025-10-31 cs.LG 75%

When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker Product

Youqi Wu, Jingwei Zhang, Farzan Farnia

机构 * Department of Computer Science & Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26146 2025-10-31 cs.LG 50%

maxVSTAR: Maximally Adaptive Vision-Guided CSI Sensing with Closed-Loop Edge Model Adaptation for Robust Human Activity Recognition

Kexing Liu

专题命中 多模态训练与对齐 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 6 篇

2308.16075 2025-10-31 cs.CL cs.AI cs.CV 82%

Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages

Baban Gain, Dibyanayan Bandyopadhyay, Samrat Mukherjee, Chandranath Adak, Asif Ekbal

机构 * Indian Institute of Technology Patna India Indian Institute of Technology Jodhpur \& Indian Institute of Technology Patna India Indian Institute of Technology Patna Indian Institute of Technology Jodhpur \& Indian Institute of Technology Patna

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14879 2025-10-31 cs.HC cs.AI 80%

Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM

Jiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube, Bingsheng Yao, Xuhai Xu, Dakuo Wang, Elizabeth Mynatt, Varun Mishra

机构 * Northeastern University(东北大学) Columbia University(哥伦比亚大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Jiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube, Bingsheng Yao, Xuhai Xu, Dakuo Wang, Elizabeth Mynatt, and Varun Mishra. 2025. Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 9, 3, Article 101 (September 2025), 37 pages

Journal ref Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 9 (2025) 101:1-37

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25811 2025-10-31 stat.ML cs.LG math.ST stat.TH 78%

Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms

William Réveillard, Richard Combes

机构 * Division of Decision and Control Systems(决策与控制系统系) KTH Royal Institute of Technology(皇家理工学院) Laboratoire des signaux et systèmes(信号与系统实验室) Université Paris-Saclay, CNRS, CentraleSupélec(巴黎萨克雷大学、国家科学研究中心、中央理工-巴黎高等理工学院)

专题命中 其他多模态 :multimodal(title,abstract)

Comments 31 pages; NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06455 2025-10-31 stat.AP 78%

Combining Unsupervised Learning and Statistical Inference For Multimodal N-of-1 Trials

Juliana Schneider, Thomas Gärtner, Stefan Konigorski

专题命中 其他多模态 :multimodal(title,abstract)

Comments 22 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14269 2025-10-31 cs.LG cs.CV cs.IR 57%

Deep Learning for Technical Document Classification

Shuo Jiang, Jie Hu, Christopher L. Magee, Jianxi Luo

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments 16 pages, 8 figures, 9 tables

Journal ref IEEE Transactions on Engineering Management 71 (2024): 1163-1179

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26209 2025-10-31 astro-ph.HE 50%

Multi-Faceted Emission Properties of PSR J2129+4119 Observed with FAST

Habtamu Menberu Tedila, Di Li, Pei Wang, Rai Yuen, Ziwei Wu, Shijun Dang, Jianping Yuan, Na Wang, Marilyn Cruces, Jun Shuo Zhang, Juntao Bai, De Zhao, FAST Collaboration

专题命中 其他多模态 :multi-modal(abstract)

Comments 22 pages, 18 figures, Revised for ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏