arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-30 至 2025-07-30 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 7 篇

2507.21395 2025-07-30 cs.MM cs.AI cs.SD eess.AS 89%

Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion

Zeyu Deng, Yanhui Lu, Jiashu Liao, Shuang Wu, Chongfeng Wei

机构 * James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院) University of Bristol(布里斯托大学) School of Engineering Mathematics and Technology, University of Bristol(布里斯托大学工程数学与技术学院) School of Computing Science, University of Glasgow(格拉斯哥大学计算科学学院) Department of Civil, Environmental & Geomatic Engineering, University College London (UCL)(伦敦大学学院(UCL)土木、环境与测绘工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08204 2025-07-30 cs.CV 83%

One-stage Modality Distillation for Incomplete Multimodal Learning

Shicai Wei, Yang Luo, Chunbo Luo

机构 * School of Information and Communication Engineering University of Electronic Science and Technology of China(信息与通信工程学院 电子科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03248 2025-07-30 cs.CV cs.AI cs.CL 82%

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Yiwu Zhong, Zhuoming Liu, Yin Li, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 79%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10774 2025-07-30 cs.LG cs.AI 79%

Context-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting

Yueyang Yao, Jiajun Li, Xingyuan Dai, MengMeng Zhang, Xiaoyan Gong, Fei-Yue Wang, Yisheng Lv

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10091 2025-07-30 eess.IV 78%

G$^{2}$SF-MIAD: Geometry-Guided Score Fusion for Multimodal Industrial Anomaly Detection

Chengyu Tao, Xuanming Cao, Juan Du

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21857 2025-07-30 cs.CV 57%

Unleashing the Power of Motion and Depth: A Selective Fusion Strategy for RGB-D Video Salient Object Detection

Jiahao He, Daerji Suolang, Keren Fu, Qijun Zhao

机构 * College of Computer Science, Sichuan University(四川大学计算机学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments submitted to TMM on 11-Jun-2024, ID: MM-020522, still in peer review

详情

展开后加载摘要…

URL PDF HTML 收藏