arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-13 至 2025-08-13 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 5 篇

2502.12454 2025-08-13 cs.CV cs.AI cs.HC cs.LG 82%

Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches

He Zhang, Xinyi Fu

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, accepted to MRAC'25: 3rd International Workshop on Multimodal and Responsible Affective Computing (ACM-MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02935 2025-08-13 cs.CL 79%

Dynamic Graph Neural ODE Network for Multi-modal Emotion Recognition in Conversation

Yuntao Shou, Tao Meng, Wei Ai, Keqin Li

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学教育部长江网络与网络安全重点实验室) College of Computer and Mathematics, Central South University of Forestry and Technology(中南林业科技大学计算机与数学学院) Department of Computer Science, State University of New York(纽约州立大学新帕尔茨分校计算机科学系)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08541 2025-08-13 physics.app-ph 78%

Multimodal learning enables instant ionizing radiation alerts on unmodified mobile phones for real-world emergency response

Yanfeng Xie, Xingzhi Cheng

专题命中 视频多模态 :multimodal(title,abstract)

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08989 2025-08-13 cs.CV 57%

KFFocus: Highlighting Keyframes for Enhanced Video Understanding

Ming Nie, Chunwei Wang, Hang Xu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08590 2025-08-13 cs.CV cs.HC 57%

QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection

Yuxiao Wang, Wolin Liang, Yu Lei, Weiying Xue, Nan Zhuang, Qi Liu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏