arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-31 至 2025-10-31 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 3 篇

2510.26027 2025-10-31 cs.CV 57%

Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders

Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk, Mohsen Fayyaz

机构 * Leibniz University Hannover(莱比锡大学汉诺威分校) L3S Research Center(L3S研究中心) Microsoft(微软公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01701 2025-10-31 cs.CV 57%

Signal-SGN: A Spiking Graph Convolutional Network for Skeletal Action Recognition via Learning Temporal-Frequency Dynamics

Naichuan Zheng, Yuchen Du, Hailun Xia, Zeyu Liang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26628 2025-10-31 cs.NI eess.SP 50%

Low-Altitude UAV-Carried Movable Antenna for Joint Wireless Power Transfer and Covert Communications

Chuang Zhang, Geng Sun, Jiahui Li, Jiacheng Wang, Qingqing Wu, Dusit Niyato, Shiwen Mao, Tony Q. S. Quek

专题命中 视频多模态 :multimodal(abstract)

Comments This paper has been submitted to IEEE Journal on Selected Areas in Communications

详情

展开后加载摘要…

URL PDF HTML 收藏